Code Speaks: Inspecting Local Government Data and Systems with an Engineer's Lens

どうも〜おかむーです! Today let's do a little GovTech post-mortem: I’ll read a manifesto in code — check how municipal datasets and systems actually hold up from a developer and data-ops point of view.
- Quick read: many Japanese municipal portals publish useful CSVs but still mix PDFs and non-machine formats.
- Main point: machine-readability, stable APIs, and metadata are the low-hanging fruit to unlock policy accountability.
- Outcome: concrete engineering steps to turn PDFs and scattered CSVs into verifiable, auditable datasets.
結論
Open data exists in Japan — Tokyo’s catalog (catalog.data.metro.tokyo.lg.jp), Saitama (opendata.pref.saitama.lg.jp), Chuo Ward (city.chuo.lg.jp) — but the bottleneck is format and discoverability. 要するに、政策の数値目標を検証するためには、PDFに埋めた目標値をCSV/JSONで公開し、APIで継続的に取得できるようにするのが最優先です。
Report: technical assessment and concrete fixes
What I inspected
These public portals publish CSVs and Excel files (good!), but many policy documents — targets, methodologies, evaluation reports — remain PDFs or HTML reports. GovTech Tokyo notes they’re consolidating dashboards (govtechtokyo.or.jp) — promising, but inconsistent across municipalities.
Problems found (engineer’s view)
- Mixed formats: CSV + PDF + Xlsx — parsing friction.
- No consistent metadata (no DCAT/JSON-LD) — hard to discover dataset semantics.
- Weak or absent APIs — scraping required for automation.
- Versioning & provenance missing — can’t trace which dataset supported a past policy claim.
これ見てくださいよ: if a policy target is in a PDF and the outcome is in a CSV, you need a pipeline to extract, normalize, and join them. That’s tedious but solvable.
Practical code snippets
Fetch a CSV and load with pandas:
import requests
import pandas as pd
r = requests.get('https://catalog.data.metro.tokyo.lg.jp/dataset/xxx.csv')
open('data.csv','wb').write(r.content)
df = pd.read_csv('data.csv')
print(df.head())
Extract tables from PDF targets (if they’re not published as CSV):
# using tabula-py
pip install tabula-py
python -m tabula --pages all --output targets.csv policy_report.pdf
Data engineering recommendations
- Publish policy targets and indicators as machine-readable JSON/CSV alongside human PDFs. Use DCAT metadata and schema.org for datasets.
- Provide a RESTful catalog API (pagination, filtering, CORS) or adopt CKAN/Socrata so external tools can query datasets reliably.
- Add dataset versioning + changelogs to enable audits.
- Standardize date formats, geocodes, and unique identifiers (e.g., j-loc codes) to ease joins.
How to measure policy gaps
まとめ
Municipal open data in Japan has great raw material — Tokyo and many prefectures already publish catalogs — but without machine-readability, API access, metadata, and versioning, you can’t reliably audit policy promises. エンジニア的に言うと、APIs + metadata + reproducible ETL pipelinesがあれば、政策の説明責任が一気に上がります!
おかむーから一言
I’ve built products on messy government data — make it machine-readable, and the world of civic innovation explodes. Let’s push for APIs, not PDFs!
Sources
- https://www.zhihu.com/question/40553450
- https://www.govtechtokyo.or.jp/services/data-utilization/
- https://www.zhihu.com/tardis/zm/art/1924492115896960699
- https://note.govtechtokyo.jp/n/n77785a8254d6
- https://www.zhihu.com/question/372341437
- https://www.zhihu.com/question/290714454
- https://metidx-gov.note.jp/n/n9468573c213b
- https://www.zhihu.com/question/6430289390
- https://zenn.dev/govtechtokyo/articles/b65dc687e50918
- https://www.zhihu.com/question/38923279
- https://catalog.data.metro.tokyo.lg.jp/dataset
- https://opendata.pref.saitama.lg.jp/
- https://www.city.niigata.lg.jp/shisei/seisaku/it/open-data/index.files/csv_manual_v1.1.pdf
- https://opendata.pref.saitama.lg.jp/datasets
- https://www.city.chuo.lg.jp/kusei/gaiyou/toukeidate/opendata.html
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.