Code-Driven Manifesto: Inspecting Japan's Gov Data and Systems

Hey — Okamu here! Today I'm taking a coder's scalpel to how Japanese national and local governments publish data and run digital services.
- Many official reports and KPI statements live in PDFs or scattered HTML, not in machine-friendly tables
- Some platforms (e-Gov, Digital Agency, e-Stat) exist, but APIs and consistent schemas are spotty
- With modest engineering fixes (CSV/JSON export, API endpoints, schema, dashboards) reuse and accountability jump dramatically
結論
Open data is present but brittle: governments publish useful numbers (see e-Gov https://e-gov.go.jp, Digital Agency https://digital.go.jp, e-Stat https://www.e-stat.go.jp), yet format and discoverability block scalable analysis. 要するに、API-firstで機械可読にすれば政策の検証と改善がぐっと楽になるということです。
Report: where things are now, technically
What I found
- Portal & services: e-Gov (policies, procedures), e-Gov electronic applications (https://shinsei.e-gov.go.jp) and Digital Agency host strategy docs and guidance PDFs.
- Local reporting: municipalities publish evaluation pages (e.g. Sukagawa city digital strategy results) often as HTML or PDF blobs.
- GovTech initiatives: GovTech Tokyo shows dashboards and reuse efforts (https://govtechtokyo.or.jp), demonstrating what's possible.
これ見てくださいよ: many KPI tables are embedded in PDFs or buried in HTML. That kills automation.
Machine-readability & APIs
- e-Stat provides a proper statistical API (useful!), but coverage varies by administrative program.
- Many ministry reports lack stable JSON/CSV endpoints; instead you get annual PDF bundles.
- Authentication and data catalogs are inconsistent between central and local governments.
エンジニア的に言うと、this is an API-design and data-engineering problem: provide stable endpoints, versioned schemas, and machine-readable releases.
Code example: pragmatic ingestion
Here's a minimal Python pattern: prefer CSV/JSON endpoints; fallback to PDF table extraction.
# prefer API/CSV
import requests, pandas as pd
url = 'https://example.gov/data/kpi.csv' # replace with real CSV endpoint
r = requests.get(url)
open('kpi.csv','wb').write(r.content)
df = pd.read_csv('kpi.csv')
print(df.head())
fallback: extract table from PDF (requires tabula-py or camelot)
pip install tabula-py
import tabula
tables = tabula.read_pdf('report.pdf', pages='all')
print(tables[0].head())
KPI gaps: targets vs reported
Example: the Digital Rural City program has published strategy docs and KPI guidance (see Digital Agency guidance), and some municipalities publish performance pages (Sukagawa). But:
- Targets are often annual milestones in prose, not as time-series values
- Results are snapshots in PDFs, so trend analysis requires manual extraction
Practical improvements (engineering roadmap)
- Catalogue-first: adopt a CKAN-style portal per prefecture / national aggregator with dataset metadata and stable URLs
- API-first distribution: every KPI timeseries available as JSON/CSV, with JSON Schema and semantic metadata (DCAT)
- CI for data: automated validation (Great Expectations) and publish-release pipelines so each dataset has lineage
- UI/UX: make CSV/JSON download buttons prominent, and expose simple REST endpoints for dashboards
- Local capacity building: templates and open-source toolkits (Dockerized ETL, simple API server) so small towns can publish structured data
まとめ
中央・地方ともにデータは出ているけど、フォーマットとAPIの欠如で使いにくい。要するに、機械可読化とAPI整備を進めれば、政策の検証・改善の速度は確実に上がるんですよね。
おかむーから一言
I've built products and startups that consume messy gov data — give me CSVs and stable APIs and I'll ship insights in a week. Technology can make government accountable and useful; let's code that manifest into reality!
Sources
- https://www.zhihu.com/question/290714454
- https://metidx-gov.note.jp/n/n9468573c213b
- https://www.zhihu.com/question/6430289390
- https://zenn.dev/govtechtokyo/articles/b65dc687e50918
- https://www.zhihu.com/question/38923279
- https://www.e-gov.go.jp/
- https://www.govtechtokyo.or.jp/services/data-utilization/
- https://shinsei.e-gov.go.jp/
- https://note.govtechtokyo.jp/n/n77785a8254d6
- https://shinsei.e-gov.go.jp/contents/preparation
- https://ja.wikipedia.org/wiki/%E3%83%87%E3%82%B8%E3%82%BF%E3%83%AB
- https://www.chisou.go.jp/sousei/pdf/r5_guideline-checkaction.pdf
- https://column.nippoukun.bpsinc.jp/what-is-digitization/
- https://www.city.sukagawa.fukushima.jp/shisei/gyoseiunei/keikaku/chiho_sosei/1015604/4045.html
- https://www.digital.go.jp/
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.