Manifesto in Code: Auditing Government Data & APIs from an Engineer’s POV

どうも〜おかむーです! Today I want to poke around government data like an engineer — friendly, practical, and a bit opinionated.
- Government publishes APIs and catalogs, but machine-readability and formats vary a lot
- PDF-first releases still block reuse; CSV/JSON + stable APIs win for civic tech
- Small engineering fixes (schema, consistent IDs, API docs) unlock huge value
結論
Public data is increasingly available (see e-Gov administrative APIs and Tokyo’s Open Data API), but the real win is treating data as code: machine-readable formats, stable endpoints, and clear schemas. PDFs and ad-hoc Excel dumps are still common and slow down reuse — fix those and civic innovation scales.
Report
What I looked at
Look, e-Gov’s Administrative API portal (e-gov.go.jp) and Tokyo’s Open Data API (portal.data.metro.tokyo.lg.jp) are doing the right thing by centralizing endpoints. The Digital Agency’s machine-readability rules (digital.go.jp) now prescribe levels (CSV/Excel/JSON as Level 1), which matters a ton.
Engineer-wise, here are the recurring issues:
- Formats: PDFs and Word docs still appear in official releases. That forces screen-scraping or OCR — costly and error-prone. 要するに、機械で読み取れないファイルが多いということです。
- Inconsistent schemas: column names change between releases, IDs are absent, timestamps use mixed timezones. For code, that's a pain.
- Thin or missing API docs: some endpoints exist but lack examples, rate limits, or stable versioning.
- Publication gaps: goal metrics in policy PDFs aren't always published as time-series data via APIs, so you can’t programmatically track progress.
Concrete technical checks
- Machine-readability rule: The Digital Agency PDF (meeting notes) explicitly requires CSV/Excel/JSON for Level 1. If a dataset is only in PDF, it fails the bar.
- API presence: Tokyo provides structured endpoints for public facilities and other resources — that’s excellent for routing data into apps.
Small code example — fetch Tokyo public facilities and write CSV
import requests, csv
r = requests.get('https://portal.data.metro.tokyo.lg.jp/opendata-api/PublicFacility')
data = r.json() # assuming JSON response
with open('facilities.csv','w',newline='') as f:
writer = csv.writer(f)
writer.writerow(['id','name','lat','lon','type'])
for item in data.get('results', []):
writer.writerow([item.get('id'), item.get('name'), item.get('lat'), item.get('lon'), item.get('type')])
要するに、API一本で多くの課題は解決できますよね。
Policy metrics vs reality
Many policy documents set numeric targets, but the data backing progress is buried in PDFs or absent from catalog portals like DATA.GO.JP or e-Stat. Engineer-wise: if you can’t pull a time series with a stable ID, you can’t build dashboards or alarms. That’s a governance gap as much as a tech one.
Practical improvement roadmap
- Mandate Level-1 machine-readable exports (CSV/JSON) for all KPI reporting — enforce via the Digital Agency rules
- Canonical schemas + semantic IDs (use URNs or gov IDs) and versioned API endpoints
- Publish OpenAPI specs for every API and host example clients
- Small middleware: automated ETL from legacy PDFs to canonical CSV with provenance metadata
まとめ
APIs and portals exist and are improving, but the bottleneck is format consistency and discoverability. Treat data releases like software releases: versioning, docs, tests, and CI for datasets. Do that and civic tech moves from ad-hoc hacks to reliable services.
おかむーから一言
I’ve built and shipped products that rely on flaky public data — believe me, invest in schema and APIs now and you’ll save months later. Let’s code the manifesto into the data itself!
Sources
- https://www.zhihu.com/question/290714454
- https://metidx-gov.note.jp/n/n9468573c213b
- https://www.zhihu.com/question/6430289390
- https://zenn.dev/govtechtokyo/articles/b65dc687e50918
- https://www.zhihu.com/question/38923279
- https://www.e-gov.go.jp/digital-government/api
- https://portal.data.metro.tokyo.lg.jp/opendata-api/
- https://api-catalog.e-gov.go.jp/info/ja/apicatalog/list
- https://japan-opendata.github.io/awesome-japan-opendata/
- https://odcs.bodik.jp/developers/
- https://www.digital.go.jp/assets/contents/node/basic_page/field_ref_resources/256dcba6-b936-4031-b88d-3abb27e27f9b/f7af0ca4/20260331_meeting_executive_outline_06.pdf
- https://www.cas.go.jp/jp/seisaku/digital_gyozaikaikaku/kakusyoDX4/kakusyoDX4.html
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.