Code-Driven Manifesto: Making Japan’s Public Data Truly Machine-Readable

IT Policy Proposals
Code-Driven Manifesto: Making Japan’s Public Data Truly Machine-Readable

Hey — おかむー here! Today I want to nerd out a bit about public data in Japan and why small tech fixes can unlock big policy wins.

  • Governments publish lots of PDFs and CSVs, but formats and metadata are inconsistent.
  • Machine-readability rules exist (e.g. Soumu guidelines) but adoption and APIs are patchy.
  • Practical fixes: CSVW/JSON-LD metadata, stable APIs, schema checks, and reproducible pipelines.

結論

Public data policy is only as good as its engineering. Japan already has rules (e.g. Ministry of Internal Affairs unified CSV guidance: https://www.soumu.go.jp/...), and pockets of CSV data (e.g. notice.go.jp CSVs), but many KPI reports and grants remain in PDFs or ad-hoc CSVs. Engineer-wise, the gap is operational: provide machine-readable contracts (OpenAPI + CSVW), validation pipelines, and publishing best-practices so policy metrics become verifiable and reusable.

Report

What I looked at

These examples show the current landscape: the Soumu unified rule (machine-readable tables) and Cabinet Office notes on AI-ready data (https://www.cas.go.jp/...), plus actual CSV endpoints like NICTER notice (https://notice.go.jp/docs/status_nicter.csv) and assorted ministry CSVs (MHLW, env, local prefecture PDFs). Look at this — some datasets are CSVs downloadable directly, some are PDFs with embedded tables (pref.yamaguchi.lg.jp reports), and many lack standardized metadata.

Technical problems observed

  • Encoding & schema drift: CSVs use different encodings (UTF-8, Shift_JIS), inconsistent headers, mixed date formats. That makes joins painful.
  • PDFs vs CSV: KPIs are published in PDFs (visual but not machine-readable). 要するに、機械が読めないってことです。
  • No API contract: no OpenAPI/OpenData portal metadata (DCAT or CSVW), so automated pipelines break on every change.
  • Missing provenance/versioning: no checksums, no dataset version tags, hard to reproduce historical analyses.

Concrete engineering fixes

  • Publish CSVW/JSON-LD metadata alongside every CSV (columns, types, encoding, sample rows). Example spec: https://www.w3.org/TR/tabular-data-primer/
  • Provide a simple REST API (OpenAPI) wrapping datasets; even a thin layer that serves CSV + schema endpoint is huge.
  • CI for data quality: run GitHub Actions that validate encodings, date formats, null rates, and report regressions.
  • For legacy PDFs: run scheduled extraction (tabula/ Camelot) with OCR fallback; export both raw and cleaned CSVs and publish diffs.

Quick code example (Python) — fetch CSV robustly

import requests

import pandas as pd

r = requests.get('https://notice.go.jp/docs/status_nicter.csv')

try utf-8 then shift_jis

for enc in ('utf-8','shift_jis','cp932'):

try:

df = pd.read_csv(pd.io.common.StringIO(r.content.decode(enc)))

break

except Exception:

continue

print(df.head())

要するに、encoding fallback と schema validation を入れておくだけでデータ再利用性が劇的に上がります!

Policy KPI gap analysis (example)

Digital rural grants (デジタル田園都市国家構想) set KPIs in program PDFs, and local reports (pref.yamaguchi.lg.jp) show project-level spend/outputs. But there’s no machine-checked aggregation across municipalities. That means the national dashboard cannot automatically verify reported KPIs — increasing audit cost and reducing trust.

Open-data potential

With standard metadata and APIs, civic techs can build monitoring dashboards, automated reconciliation with financial systems, and alerting for KPI underperformance. That’s low-hanging fruit for transparency and AI-ready government data (as the Cabinet Office notes).

まとめ

Small engineering patterns — CSVW metadata, OpenAPI wrappers, encoding checks, PDF extraction pipelines, and CI data quality tests — unlock massive policy value. The rules exist; now it’s about operationalizing them across ministries and local governments.

おかむーから一言

I’ve built and shipped data products in startups and GovTech — the tech is straightforward, the challenge is coordination. Let’s put simple machine-readable contracts around datasets and treat data publishing like product delivery. Tech + policy = multiplier effect!