Code-Driven Manifest: Making Government Data Actually Useful

IT Policy Proposals
Code-Driven Manifest: Making Government Data Actually Useful

どうも〜おかむーです! Hi — Okamu here, and today we're doing a bit of GovTech autopsy with an engineer's scalpel.

  • Governments publish a lot, but much is stuck in PDFs and proprietary systems.
  • Machine-readability rules (digital.go.jp, soumu.go.jp) are improving, but gaps remain in APIs, metadata, and formats.
  • Technical fixes (CSV/JSON, OpenAPI, DCAT, automated pipelines) can close the gap fast and unlock reuse.

結論

Public-sector datasets are moving in the right direction — see the Digital Agency and Soumu initiatives — but practical usability still falters because many statistical tables and administrative documents remain trapped in PDF/Word or inconsistent formats. Engineer-wise, this is mostly an interoperability and automation problem: publish standard machine formats (CSV/JSON/JSON-LD), deliver stable APIs with OpenAPI specs, and add rich metadata (DCAT/Schema.org). That unlocks analytics, dashboards, and civic apps quickly.

Report

What I looked at

  • Machine-readability guidance draft from Digital Agency (digital.go.jp, meeting doc 2026-03-31) emphasizes Excel/CSV over PDF for Level 1 datasets.
  • Soumu's unified rules for machine-readable statistical tables (soumu.go.jp, 2020) set practical conventions for table structure and metadata.
  • Local examples: Fukuoka public-facility reservation pages (city.fukuoka.lg.jp / user-facing portal) show operational services but limited open API endpoints.

これ見てくださいよ — the recurring pattern is clear: policy and guidance exist, but implementation varies across municipalities.

Technical issues observed

  • Format fragmentation: many releases are PDFs or XLSX with inconsistent headers and merged cells. That breaks automated ingestion.
  • Missing machine endpoints: few municipalities provide RESTful APIs with schema docs. Where APIs exist, authentication and rate limits are inconsistent.
  • Sparse metadata: update frequency, provenance, and license are often absent or buried in PDFs, making trust and reuse harder.
  • Policy vs reality gap: targets for cloud migration and standardization (Digital Agency roadmap toward 2025/2026) are ambitious, but operational metrics (how many datasets are machine-readable and API-backed) are not centrally published.

Concrete engineering fixes (practical)

  • Publish a canonical CSV/JSON alongside any PDF. Use stable filenames and include a manifest.json with dataset metadata.
  • Adopt DCAT-AP / schema.org/dataset for catalog metadata so search engines and civic apps discover datasets.
  • Provide a lightweight OpenAPI spec for any API endpoints. Example minimal OpenAPI approach: a single /datasets endpoint returning paginated JSON.

Example ingest snippet (Python/pandas) — how an app would consume a CSV endpoint:

import requests

import pandas as pd

url = 'https://example.city.gov/datasets/facilities.csv'

resp = requests.get(url)

resp.raise_for_status()

df = pd.read_csv(pd.compat.StringIO(resp.text))

normalize columns

df.columns = [c.strip().lower().replace(' ', '_') for c in df.columns]

print(df.head())

Note: the snippet is illustrative — it shows how straightforward consumption becomes when CSV/JSON endpoints exist.

Data quality & KPI alignment

  • Policy targets should be measurable: e.g., percent of core administrative datasets available as CSV/JSON and percent with OpenAPI. Track these in a public dashboard.
  • Example KPI: "By Q4 2026, 80% of municipal facility and budget tables must be machine-readable and documented in a catalog." Then audit monthly using automated checks (schema, sample rows, license presence).

Reuse opportunities

  • Real-time reservation availability (Fukuoka-style systems) + open facilities data => civic apps for booking optimization.
  • Standardized budget line-items as JSON => comparative analytics and watchdog dashboards.

まとめ

Policy momentum is real: Digital Agency and Soumu guidance gives us the rules. The next step is operational: standardize formats (CSV/JSON/JSON-LD), publish machine endpoints with OpenAPI, and add catalog metadata (DCAT/schema). That's a short engineering path from PDFs to reusable civic data.

おかむーから一言

Tech can make government faster and fairer — but only if we stop treating data like printed paper. Let's ship APIs, not PDFs. I'm ready to build the pipelines with anyone who'll listen!