Code-Driven Manifesto: Making Japan's Public Data Truly Machine-Readable

どうも〜おかむーです! Today I'm taking a coder's lens to public-sector data in Japan — the place where PDFs go to retire and APIs get ghosted. エンジニア的に言うと、データが使えないと政策も検証できないんですよね。
- Local and central gov are moving toward machine-readable rules, but many datasets are still PDFs.
- The Digital Agency and Cabinet Secretariat have put concrete rules and decisions on machine readability — implementation gaps remain.
- Technical fixes are straightforward: standard schema, a central API gateway, and sane metadata + tooling.
結論
Japan's policy direction is correct: require CSV/Excel/API-first publishing (see Digital Agency drafts at digital.go.jp and Cabinet Secretariat materials at cas.go.jp). However, the reality across municipalities is mixed—lots of PDF-locked reports and inconsistent schemas. 要するに、ルールはあるけど実装が追いついてないということです。エンジニア的には、API一本で解決する話なんですよ。
What I examined
Sources: Digital Agency's machine-readability rules draft (digital.go.jp), Cabinet Secretariat's digital governance pages (cas.go.jp), and GovTech Tokyo case studies showing practical reuse (govtechtokyo.or.jp / note.govtechtokyo.jp).
Machine-readability: PDF vs CSV
これ見てくださいよ:政府資料はまだPDFで公開されていることが多い。Digital Agency's draft explicitly ranks formats (Level 1: viewable/transcribable, Level 2: machine readable like CSV/Excel) — so PDF-only is considered insufficient. 要するに、PDFは人間には読めてもコードには読めないってことです。
技術的問題点:
- Parsing PDFs is brittle (tables span pages, encodings, merged cells).
- No standardized field names across municipalities ("population" vs "pop" vs "人数").
- Missing machine-readable metadata (provenance, update frequency, license).
API and format availability
Many gov datasets lack a consistent REST API or OpenAPI description. Where APIs exist, schemas differ and authentication approaches vary. From an engineering perspective, that increases integration cost exponentially.
Quick code example: measure municipal compliance rate
import requests, pandas as pd
Hypothetical catalog endpoint returning JSON metadata
catalog = requests.get('https://example-gov.jp/api/datasets').json()
rows = []
for ds in catalog['datasets']:
fmt = ds.get('format','').lower()
rows.append({'id': ds['id'], 'format': fmt})
df = pd.DataFrame(rows)
rate = (df['format'].isin(['csv','xlsx'])).mean()
print(f"Machine-readable rate: {rate:.0%}")
要するに、まずはカタログからフォーマットを数えるだけで現状把握できます。
PDF rescue: pragmatic ETL patterns
- Use table-extraction tools (tabula-py / camelot) as a stopgap.
- Prefer automated OCR + heuristic column mapping, but log manual review.
- Store canonical CSV in a versioned S3 bucket and publish an API layer on top.
Example transform pipeline:
- Ingest: fetch PDF from lg.jp site
- Extract: camelot -> raw CSV
- Normalize: map column names to standard schema (use mapping table)
- Validate: run schema checks (great_expectations)
- Publish: push to API endpoint + metadata registry (DCAT-compliant)
Policy targets vs reality
The Cabinet Secretariat and Digital Agency have set rules and meeting outputs to improve machine-readability, but without enforcement timelines many municipalities lag. So numeric targets (percentage of datasets machine-readable) should be tracked via a central dashboard — GovTech Tokyo's dashboards are a good template.
Improvements I recommend (concrete)
- Mandate an open catalog (DCAT) under a gov-central endpoint (go.jp). Include format, schema link, license, and update cadence.
- Require OpenAPI/Swagger for any dataset API and publish examples.
- Provide a reference Python/Node SDK and CI templates so municipalities can publish with one click.
- Deploy a central ETL toolkit (containerized) that municipalities can reuse to convert PDFs to canonical CSV + tests.
- Measure and publish a monthly "machine-readable rate" on a public dashboard (e-Stat style).
まとめ
Policy rules are landing — Digital Agency and Cabinet Secretariat materials show momentum — but the hard work is operational: standard schemas, central catalog, APIs, and reusable tooling. エンジニア的に言うと、ここは改善余地ありまくり!でも技術的対処法は明確で、やればかなり成果が出ます。
おかむーから一言
I've built products and startups that ship under messy constraints — this is solvable. Tech + policy aligned = measurable public value. Let's make government data actually usable, one API at a time!
Sources
- https://ja.wikipedia.org/wiki/%E5%85%AC%E5%85%B1
- https://www.intec.co.jp/column/smartcity-08.html
- https://kotobank.jp/word/%E5%85%AC%E5%85%B1-494676
- https://www.digital.go.jp/resources/data_case_study_private
- https://www.takeda.tv/saga/blog/post-230907/
- https://www.zhihu.com/question/290714454
- https://www.digital.go.jp/assets/contents/node/basic_page/field_ref_resources/256dcba6-b936-4031-b88d-3abb27e27f9b/f7af0ca4/20260331_meeting_executive_outline_06.pdf
- https://www.zhihu.com/question/6430289390
- https://www.cas.go.jp/jp/seisaku/digital_gyozaikaikaku/kakusyoDX4/kakusyoDX4.html
- https://www.zhihu.com/question/38923279
- https://www.zhihu.com/question/40553450
- https://www.govtechtokyo.or.jp/services/data-utilization/
- https://www.zhihu.com/tardis/zm/art/1924492115896960699
- https://note.govtechtokyo.jp/n/n77785a8254d6
- https://www.zhihu.com/question/372341437
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.