Code the Manifesto: Auditing Japan's Local Government Data Stack

どうも〜おかむーです! Hi — today we're doing a slightly nerdy, slightly political audit of how Japan's government and local municipalities publish data. エンジニア的に言うと、政策はコードとデータで検証できるんですよ〜
- The national push mandates standardized core systems for local governments, but reality is fragmented data formats and legacy systems.
- Many useful datasets exist (Japan Dashboard, e-Stat, individual CSVs), yet machine-readability and APIs are inconsistent.
- Practical fixes: schema-first APIs, a central data catalog, CI for data quality, and migration tooling for PDF→CSV/JSON.
結論
Digital Agency's law and the Japan Dashboard are the right signals, but implementation gaps remain. 要するに、ルールはあるけど現場でデータが使いやすくなっていないということです。標準化は進めつつ、開発者が本当に使えるAPI・スキーマを作ることがもっと重要です。
Report: what I looked at and what it means
Sources I checked
- Digital Agency policy pages on local government core system standardization: https://www.digital.go.jp/policies/local_governments
- The local government information systems standardization law: https://laws.e-gov.go.jp/law/503AC0000000040
- Japan Dashboard landing: https://www.digital.go.jp/resources/japandashboard
- e-Stat / statistics dashboard: https://dashboard.e-stat.go.jp/
- Example direct CSV links found in search results (NICTER notice, ministries' CSV links like https://notice.go.jp/docs/status_nicter.csv, https://www.jinji.go.jp/content/900024615.csv, etc.)
これ見てくださいよ:official CSVs exist, but they are scattered across different domains, sometimes lacking metadata, and occasionally still published only as PDF.
Policy vs Implementation
- Policy: the law defines ~20 standardized business tasks and requires conformity of local systems to standardization criteria. Good — that gives a clear scope.
- Reality: many municipalities still run legacy core systems, exchange PDF reports, or publish CSVs with inconsistent headers/encodings. That blocks reuse and automation.
要するに、法律は "what" を示しているが、"how" が足りないんです。
Technical problems observed
- Machine readability: PDFs are still common; CSVs often lack schema, column types, stable IDs, timezones.
- API posture: e-Stat offers APIs, Japan Dashboard aggregates statistics, but local governments rarely expose uniform REST/GraphQL endpoints.
- Provenance & metadata: many CSV files lack machine-readable metadata (title, update frequency, licenses).
- Data quality: inconsistent encodings (Shift_JIS vs UTF-8), mixed date formats, missing primary keys.
Concrete code-centric fixes
- Schema-first approach: publish JSON Schema / OpenAPI for every dataset. This makes validation, docs, and client generation trivial.
- Central catalog: a searchable registry (harvests dataset URLs, schema, license, last-updated) — could be integrated into Japan Dashboard.
- CI/CD for data: run automated checks on CSV/JSON (schema validation, encoding, unique key checks) before publishing.
- PDF → structured data migration: use pipeline (ocr -> table extraction -> schema mapping). But better: discourage PDFs and provide source CSV/JSON.
Example: simple Python snippet to robustly read a CSV from a ministry and normalize types
import pandas as pd
url = 'https://notice.go.jp/docs/status_nicter.csv'
df = pd.read_csv(url, encoding='utf-8', parse_dates=['timestamp'], dtype={'id': str})
normalize column names
df.columns = df.columns.str.strip().str.lower().str.replace(' ', '_')
validate required columns
required = {'id','timestamp','status'}
missing = required - set(df.columns)
if missing:
raise SystemExit(f"Missing columns: {missing}")
Migration path for municipalities
- Step 1: inventory datasets and publish a minimal data catalog (CSV + simple metadata JSON).
- Step 2: add OpenAPI/JSON Schema and enable a simple REST endpoint per dataset. Use serverless functions to front legacy DBs when rewriting systems is too costly.
- Step 3: enforce data quality gates in deployment pipelines, and publish example client code.
まとめ
Japan has the right high-level policy (Digital Agency law + Japan Dashboard) and pockets of machine-friendly data (e-Stat, direct CSVs). But too much friction remains: inconsistent formats, missing metadata, and legacy systems. エンジニア的に言うと、API一本、スキーマ一つで解決する話が多いんですよね。標準はあるけど、まずは“使えるデータ”を現場に届ける実装力が必要です。
おかむーから一言
Tech can make government measurable and improvable — but only if data is actually usable. Let's ship schemas, not PDFs. I'm ready to help build the pipelines!
Sources
- https://www.digital.go.jp/policies/local_governments
- https://laws.e-gov.go.jp/law/503AC0000000040
- https://www.keiba.go.jp/
- https://www.keiba.go.jp/KeibaWeb/TodayRaceInfo/TodayRaceInfoTop
- https://www.keiba.go.jp/live/
- https://notice.go.jp/docs/status_nicter.csv
- https://www.jinji.go.jp/content/900024615.csv
- https://www.env.go.jp/content/900398071.csv
- https://www.inpit.go.jp/content/100869372.csv
- https://www.mhlw.go.jp/content/001429362.csv
- https://www.digital.go.jp/resources/japandashboard
- https://dashboard.e-stat.go.jp/
- https://www.kantei.go.jp/
- https://www.stat.go.jp/info/guide/public/kouhou/index.html
- https://www.kantei.go.jp/jp/kakugikettei/index.html
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.