Manifesto in Code: Auditing Government Data & APIs from an Engineer’s POV

IT Policy Proposals
Manifesto in Code: Auditing Government Data & APIs from an Engineer’s POV

どうも〜おかむーです! Today I want to poke around government data like an engineer — friendly, practical, and a bit opinionated.

  • Government publishes APIs and catalogs, but machine-readability and formats vary a lot
  • PDF-first releases still block reuse; CSV/JSON + stable APIs win for civic tech
  • Small engineering fixes (schema, consistent IDs, API docs) unlock huge value

結論

Public data is increasingly available (see e-Gov administrative APIs and Tokyo’s Open Data API), but the real win is treating data as code: machine-readable formats, stable endpoints, and clear schemas. PDFs and ad-hoc Excel dumps are still common and slow down reuse — fix those and civic innovation scales.

Report

What I looked at

Look, e-Gov’s Administrative API portal (e-gov.go.jp) and Tokyo’s Open Data API (portal.data.metro.tokyo.lg.jp) are doing the right thing by centralizing endpoints. The Digital Agency’s machine-readability rules (digital.go.jp) now prescribe levels (CSV/Excel/JSON as Level 1), which matters a ton.

Engineer-wise, here are the recurring issues:

  • Formats: PDFs and Word docs still appear in official releases. That forces screen-scraping or OCR — costly and error-prone. 要するに、機械で読み取れないファイルが多いということです。
  • Inconsistent schemas: column names change between releases, IDs are absent, timestamps use mixed timezones. For code, that's a pain.
  • Thin or missing API docs: some endpoints exist but lack examples, rate limits, or stable versioning.
  • Publication gaps: goal metrics in policy PDFs aren't always published as time-series data via APIs, so you can’t programmatically track progress.

Concrete technical checks

  • Machine-readability rule: The Digital Agency PDF (meeting notes) explicitly requires CSV/Excel/JSON for Level 1. If a dataset is only in PDF, it fails the bar.
  • API presence: Tokyo provides structured endpoints for public facilities and other resources — that’s excellent for routing data into apps.

Small code example — fetch Tokyo public facilities and write CSV

import requests, csv

r = requests.get('https://portal.data.metro.tokyo.lg.jp/opendata-api/PublicFacility')

data = r.json() # assuming JSON response

with open('facilities.csv','w',newline='') as f:

writer = csv.writer(f)

writer.writerow(['id','name','lat','lon','type'])

for item in data.get('results', []):

writer.writerow([item.get('id'), item.get('name'), item.get('lat'), item.get('lon'), item.get('type')])

要するに、API一本で多くの課題は解決できますよね。

Policy metrics vs reality

Many policy documents set numeric targets, but the data backing progress is buried in PDFs or absent from catalog portals like DATA.GO.JP or e-Stat. Engineer-wise: if you can’t pull a time series with a stable ID, you can’t build dashboards or alarms. That’s a governance gap as much as a tech one.

Practical improvement roadmap

  • Mandate Level-1 machine-readable exports (CSV/JSON) for all KPI reporting — enforce via the Digital Agency rules
  • Canonical schemas + semantic IDs (use URNs or gov IDs) and versioned API endpoints
  • Publish OpenAPI specs for every API and host example clients
  • Small middleware: automated ETL from legacy PDFs to canonical CSV with provenance metadata

まとめ

APIs and portals exist and are improving, but the bottleneck is format consistency and discoverability. Treat data releases like software releases: versioning, docs, tests, and CI for datasets. Do that and civic tech moves from ad-hoc hacks to reliable services.

おかむーから一言

I’ve built and shipped products that rely on flaky public data — believe me, invest in schema and APIs now and you’ll save months later. Let’s code the manifesto into the data itself!