Code the Manifesto: Auditing Japan's Local Government Data Stack

IT Policy Proposals
Code the Manifesto: Auditing Japan's Local Government Data Stack

どうも〜おかむーです! Hi — today we're doing a slightly nerdy, slightly political audit of how Japan's government and local municipalities publish data. エンジニア的に言うと、政策はコードとデータで検証できるんですよ〜

  • The national push mandates standardized core systems for local governments, but reality is fragmented data formats and legacy systems.
  • Many useful datasets exist (Japan Dashboard, e-Stat, individual CSVs), yet machine-readability and APIs are inconsistent.
  • Practical fixes: schema-first APIs, a central data catalog, CI for data quality, and migration tooling for PDF→CSV/JSON.

結論

Digital Agency's law and the Japan Dashboard are the right signals, but implementation gaps remain. 要するに、ルールはあるけど現場でデータが使いやすくなっていないということです。標準化は進めつつ、開発者が本当に使えるAPI・スキーマを作ることがもっと重要です。

Report: what I looked at and what it means

Sources I checked

  • Digital Agency policy pages on local government core system standardization: https://www.digital.go.jp/policies/local_governments
  • The local government information systems standardization law: https://laws.e-gov.go.jp/law/503AC0000000040
  • Japan Dashboard landing: https://www.digital.go.jp/resources/japandashboard
  • e-Stat / statistics dashboard: https://dashboard.e-stat.go.jp/
  • Example direct CSV links found in search results (NICTER notice, ministries' CSV links like https://notice.go.jp/docs/status_nicter.csv, https://www.jinji.go.jp/content/900024615.csv, etc.)

これ見てくださいよ:official CSVs exist, but they are scattered across different domains, sometimes lacking metadata, and occasionally still published only as PDF.

Policy vs Implementation

  • Policy: the law defines ~20 standardized business tasks and requires conformity of local systems to standardization criteria. Good — that gives a clear scope.
  • Reality: many municipalities still run legacy core systems, exchange PDF reports, or publish CSVs with inconsistent headers/encodings. That blocks reuse and automation.

要するに、法律は "what" を示しているが、"how" が足りないんです。

Technical problems observed

  • Machine readability: PDFs are still common; CSVs often lack schema, column types, stable IDs, timezones.
  • API posture: e-Stat offers APIs, Japan Dashboard aggregates statistics, but local governments rarely expose uniform REST/GraphQL endpoints.
  • Provenance & metadata: many CSV files lack machine-readable metadata (title, update frequency, licenses).
  • Data quality: inconsistent encodings (Shift_JIS vs UTF-8), mixed date formats, missing primary keys.

Concrete code-centric fixes

  • Schema-first approach: publish JSON Schema / OpenAPI for every dataset. This makes validation, docs, and client generation trivial.
  • Central catalog: a searchable registry (harvests dataset URLs, schema, license, last-updated) — could be integrated into Japan Dashboard.
  • CI/CD for data: run automated checks on CSV/JSON (schema validation, encoding, unique key checks) before publishing.
  • PDF → structured data migration: use pipeline (ocr -> table extraction -> schema mapping). But better: discourage PDFs and provide source CSV/JSON.

Example: simple Python snippet to robustly read a CSV from a ministry and normalize types

import pandas as pd

url = 'https://notice.go.jp/docs/status_nicter.csv'

df = pd.read_csv(url, encoding='utf-8', parse_dates=['timestamp'], dtype={'id': str})

normalize column names

df.columns = df.columns.str.strip().str.lower().str.replace(' ', '_')

validate required columns

required = {'id','timestamp','status'}

missing = required - set(df.columns)

if missing:

raise SystemExit(f"Missing columns: {missing}")

Migration path for municipalities

  • Step 1: inventory datasets and publish a minimal data catalog (CSV + simple metadata JSON).
  • Step 2: add OpenAPI/JSON Schema and enable a simple REST endpoint per dataset. Use serverless functions to front legacy DBs when rewriting systems is too costly.
  • Step 3: enforce data quality gates in deployment pipelines, and publish example client code.

まとめ

Japan has the right high-level policy (Digital Agency law + Japan Dashboard) and pockets of machine-friendly data (e-Stat, direct CSVs). But too much friction remains: inconsistent formats, missing metadata, and legacy systems. エンジニア的に言うと、API一本、スキーマ一つで解決する話が多いんですよね。標準はあるけど、まずは“使えるデータ”を現場に届ける実装力が必要です。

おかむーから一言

Tech can make government measurable and improvable — but only if data is actually usable. Let's ship schemas, not PDFs. I'm ready to help build the pipelines!