Code-Driven Manifesto: Making Japan’s Public Data Truly Machine-Readable

Hey — おかむー here! Today I want to nerd out a bit about public data in Japan and why small tech fixes can unlock big policy wins.
- Governments publish lots of PDFs and CSVs, but formats and metadata are inconsistent.
- Machine-readability rules exist (e.g. Soumu guidelines) but adoption and APIs are patchy.
- Practical fixes: CSVW/JSON-LD metadata, stable APIs, schema checks, and reproducible pipelines.
結論
Public data policy is only as good as its engineering. Japan already has rules (e.g. Ministry of Internal Affairs unified CSV guidance: https://www.soumu.go.jp/...), and pockets of CSV data (e.g. notice.go.jp CSVs), but many KPI reports and grants remain in PDFs or ad-hoc CSVs. Engineer-wise, the gap is operational: provide machine-readable contracts (OpenAPI + CSVW), validation pipelines, and publishing best-practices so policy metrics become verifiable and reusable.
Report
What I looked at
These examples show the current landscape: the Soumu unified rule (machine-readable tables) and Cabinet Office notes on AI-ready data (https://www.cas.go.jp/...), plus actual CSV endpoints like NICTER notice (https://notice.go.jp/docs/status_nicter.csv) and assorted ministry CSVs (MHLW, env, local prefecture PDFs). Look at this — some datasets are CSVs downloadable directly, some are PDFs with embedded tables (pref.yamaguchi.lg.jp reports), and many lack standardized metadata.
Technical problems observed
- Encoding & schema drift: CSVs use different encodings (UTF-8, Shift_JIS), inconsistent headers, mixed date formats. That makes joins painful.
- PDFs vs CSV: KPIs are published in PDFs (visual but not machine-readable). 要するに、機械が読めないってことです。
- No API contract: no OpenAPI/OpenData portal metadata (DCAT or CSVW), so automated pipelines break on every change.
- Missing provenance/versioning: no checksums, no dataset version tags, hard to reproduce historical analyses.
Concrete engineering fixes
- Publish CSVW/JSON-LD metadata alongside every CSV (columns, types, encoding, sample rows). Example spec: https://www.w3.org/TR/tabular-data-primer/
- Provide a simple REST API (OpenAPI) wrapping datasets; even a thin layer that serves CSV + schema endpoint is huge.
- CI for data quality: run GitHub Actions that validate encodings, date formats, null rates, and report regressions.
- For legacy PDFs: run scheduled extraction (tabula/ Camelot) with OCR fallback; export both raw and cleaned CSVs and publish diffs.
Quick code example (Python) — fetch CSV robustly
import requests
import pandas as pd
r = requests.get('https://notice.go.jp/docs/status_nicter.csv')
try utf-8 then shift_jis
for enc in ('utf-8','shift_jis','cp932'):
try:
df = pd.read_csv(pd.io.common.StringIO(r.content.decode(enc)))
break
except Exception:
continue
print(df.head())
要するに、encoding fallback と schema validation を入れておくだけでデータ再利用性が劇的に上がります!
Policy KPI gap analysis (example)
Digital rural grants (デジタル田園都市国家構想) set KPIs in program PDFs, and local reports (pref.yamaguchi.lg.jp) show project-level spend/outputs. But there’s no machine-checked aggregation across municipalities. That means the national dashboard cannot automatically verify reported KPIs — increasing audit cost and reducing trust.
Open-data potential
With standard metadata and APIs, civic techs can build monitoring dashboards, automated reconciliation with financial systems, and alerting for KPI underperformance. That’s low-hanging fruit for transparency and AI-ready government data (as the Cabinet Office notes).
まとめ
Small engineering patterns — CSVW metadata, OpenAPI wrappers, encoding checks, PDF extraction pipelines, and CI data quality tests — unlock massive policy value. The rules exist; now it’s about operationalizing them across ministries and local governments.
おかむーから一言
I’ve built and shipped data products in startups and GovTech — the tech is straightforward, the challenge is coordination. Let’s put simple machine-readable contracts around datasets and treat data publishing like product delivery. Tech + policy = multiplier effect!
Sources
- https://www.zhihu.com/question/290714454
- https://www.soumu.go.jp/menu_news/s-news/01toukatsu01_02000186.html
- https://www.zhihu.com/question/6430289390
- https://www.cas.go.jp/jp/seisaku/digital_gyozaikaikaku/data8/data8_siryou1.pdf
- https://www.zhihu.com/question/38923279
- https://notice.go.jp/docs/status_nicter.csv
- https://www.jinji.go.jp/content/900024615.csv
- https://www.env.go.jp/content/900398071.csv
- https://www.inpit.go.jp/content/100869372.csv
- https://www.mhlw.go.jp/content/001429362.csv
- https://ja.wikipedia.org/wiki/%E3%83%87%E3%82%B8%E3%82%BF%E3%83%AB
- https://www.chisou.go.jp/sousei/pdf/r5_guideline-checkaction.pdf
- https://www.digital.go.jp/
- https://www.pref.yamaguchi.lg.jp/uploaded/attachment/160746.pdf
- https://biz.kddi.com/content/column/smartwork/what-is-digital/
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.