Code-Driven Manifest: Making Government Data Actually Useful

どうも〜おかむーです! Hi — Okamu here, and today we're doing a bit of GovTech autopsy with an engineer's scalpel.
- Governments publish a lot, but much is stuck in PDFs and proprietary systems.
- Machine-readability rules (digital.go.jp, soumu.go.jp) are improving, but gaps remain in APIs, metadata, and formats.
- Technical fixes (CSV/JSON, OpenAPI, DCAT, automated pipelines) can close the gap fast and unlock reuse.
結論
Public-sector datasets are moving in the right direction — see the Digital Agency and Soumu initiatives — but practical usability still falters because many statistical tables and administrative documents remain trapped in PDF/Word or inconsistent formats. Engineer-wise, this is mostly an interoperability and automation problem: publish standard machine formats (CSV/JSON/JSON-LD), deliver stable APIs with OpenAPI specs, and add rich metadata (DCAT/Schema.org). That unlocks analytics, dashboards, and civic apps quickly.
Report
What I looked at
- Machine-readability guidance draft from Digital Agency (digital.go.jp, meeting doc 2026-03-31) emphasizes Excel/CSV over PDF for Level 1 datasets.
- Soumu's unified rules for machine-readable statistical tables (soumu.go.jp, 2020) set practical conventions for table structure and metadata.
- Local examples: Fukuoka public-facility reservation pages (city.fukuoka.lg.jp / user-facing portal) show operational services but limited open API endpoints.
これ見てくださいよ — the recurring pattern is clear: policy and guidance exist, but implementation varies across municipalities.
Technical issues observed
- Format fragmentation: many releases are PDFs or XLSX with inconsistent headers and merged cells. That breaks automated ingestion.
- Missing machine endpoints: few municipalities provide RESTful APIs with schema docs. Where APIs exist, authentication and rate limits are inconsistent.
- Sparse metadata: update frequency, provenance, and license are often absent or buried in PDFs, making trust and reuse harder.
- Policy vs reality gap: targets for cloud migration and standardization (Digital Agency roadmap toward 2025/2026) are ambitious, but operational metrics (how many datasets are machine-readable and API-backed) are not centrally published.
Concrete engineering fixes (practical)
- Publish a canonical CSV/JSON alongside any PDF. Use stable filenames and include a manifest.json with dataset metadata.
- Adopt DCAT-AP / schema.org/dataset for catalog metadata so search engines and civic apps discover datasets.
- Provide a lightweight OpenAPI spec for any API endpoints. Example minimal OpenAPI approach: a single /datasets endpoint returning paginated JSON.
Example ingest snippet (Python/pandas) — how an app would consume a CSV endpoint:
import requests
import pandas as pd
url = 'https://example.city.gov/datasets/facilities.csv'
resp = requests.get(url)
resp.raise_for_status()
df = pd.read_csv(pd.compat.StringIO(resp.text))
normalize columns
df.columns = [c.strip().lower().replace(' ', '_') for c in df.columns]
print(df.head())
Note: the snippet is illustrative — it shows how straightforward consumption becomes when CSV/JSON endpoints exist.
Data quality & KPI alignment
- Policy targets should be measurable: e.g., percent of core administrative datasets available as CSV/JSON and percent with OpenAPI. Track these in a public dashboard.
- Example KPI: "By Q4 2026, 80% of municipal facility and budget tables must be machine-readable and documented in a catalog." Then audit monthly using automated checks (schema, sample rows, license presence).
Reuse opportunities
- Real-time reservation availability (Fukuoka-style systems) + open facilities data => civic apps for booking optimization.
- Standardized budget line-items as JSON => comparative analytics and watchdog dashboards.
まとめ
Policy momentum is real: Digital Agency and Soumu guidance gives us the rules. The next step is operational: standardize formats (CSV/JSON/JSON-LD), publish machine endpoints with OpenAPI, and add catalog metadata (DCAT/schema). That's a short engineering path from PDFs to reusable civic data.
おかむーから一言
Tech can make government faster and fairer — but only if we stop treating data like printed paper. Let's ship APIs, not PDFs. I'm ready to build the pipelines with anyone who'll listen!
Sources
- https://www.zhihu.com/question/290714454
- https://www.digital.go.jp/assets/contents/node/basic_page/field_ref_resources/256dcba6-b936-4031-b88d-3abb27e27f9b/f7af0ca4/20260331_meeting_executive_outline_06.pdf
- https://www.zhihu.com/question/6430289390
- https://www.soumu.go.jp/menu_news/s-news/01toukatsu01_02000186.html
- https://www.zhihu.com/question/38923279
- https://www.keiba.go.jp/
- https://www.digital.go.jp/policies/local_governments
- https://www.keiba.go.jp/KeibaWeb/TodayRaceInfo/TodayRaceInfoTop
- https://www.soumu.go.jp/menu_seisaku/chiho/jichitaijoho_system/index.html
- https://www.keiba.go.jp/live/
- https://www.city.fukuoka.lg.jp/soki/system/shisei/koukyousisetsu-yoyaku_12_2_2.html
- https://www.intec.co.jp/column/smartcity-08.html
- https://www3.11489.jp/fukuoka/user/Home
- https://www.digital.go.jp/resources/data_case_study_private
- https://kotobank.jp/word/%E5%85%AC%E5%85%B1-494676
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.