Code-First Manifest: How Japan's Public Data Stacks Up (and How to Fix It)

IT Policy Proposals
Code-First Manifest: How Japan's Public Data Stacks Up (and How to Fix It)

どうも〜おかむーです! Today I want to take a code-first look at how Japanese national and local governments publish data and run projects. エンジニア的に言うと、政策は仕様書よりもデータとAPIで語られるべきなんですよ〜

  • Government often publishes PDFs while some ministries offer CSVs — inconsistent machine-readability
  • Some datasets (notice.go.jp, env.go.jp CSVs) are great for engineering; others (policy PDFs like chisou.go.jp / prefecture PDFs) need heavy parsing
  • Technical fixes: standard metadata (DCAT), UTF-8 CSVs, OpenAPI endpoints, and CI for KPI tracking

結論

Public-sector data availability in Japan is improving but remains patchy: pockets of good machine-readable publishing exist (e.g., notice.go.jp CSVs, env.go.jp CSVs), while many important policy documents and KPI reports are buried in PDFs (chisou.go.jp guidelines, prefectural PDFs). 要するに、政策の実績と目標をちゃんとつなげるためには、APIと良質なCSVが必須ということです。

Report: what I inspected and what it tells us

What I looked at

  • Digital Agency portal (digital.go.jp) — strategy and coordination, but not a complete API catalog
  • Notice CSV (notice.go.jp/docs/status_nicter.csv) — example of good machine-readable publishing
  • Ministry/public CSV files (jinji.go.jp, env.go.jp, mhlw.go.jp in search results) — raw CSVs available
  • Policy/KPI PDFs (chisou.go.jp guideline-checkaction.pdf, prefectural PDF for Yamaguchi) — important metrics inside PDFs

これ見てくださいよ: notice.go.jp provides direct CSV links that are trivially ingestible. By contrast, the Digital田園都市交付金 KPI report (chisou.go.jp PDF) lists KPI values but as PDF pages — automated monitoring becomes painful.

Technical issues observed

  • Format fragmentation: CSV, PDF, and sometimes HTML tables; inconsistent schemas
  • Encoding problems: many legacy CSVs may be Shift_JIS; need UTF-8 normalization
  • Lack of machine-readable metadata: no DCAT or machine-readable schema to describe datasets
  • API gaps: few datasets expose OpenAPI or REST endpoints; teams resort to scraping PDFs
  • KPI publication practice: targets vs results are published but not in time-series-friendly formats

Concrete technical checks and patterns

  • Prefer direct CSV/JSON endpoints over PDFs. If only PDF is available, use pdfplumber/tabula for extraction, then validate with pandas
  • Example Python snippet (conceptual):
import pandas as pd

url = 'https://www.jinji.go.jp/content/900024615.csv'

df = pd.read_csv(url, encoding='shift_jis')

print(df.head())

  • For PDFs: use tabula-py to extract tables, then normalize column names and date formats (ISO 8601). 要するに、パースと正規化が必要ってことです。

KPI gap analysis (methodology)

  • Pull KPI tables from PDF or CSV, normalize to time-series, and compare target vs actual
  • Flag datasets where actuals are only available as yearly PDF snapshots — then set up automated extraction + unit tests
  • Example: Yamaguchi prefecture's report shows allocated amounts (e.g., 15,289千円) inside a PDF. If you want to trend funding vs outcomes, you need that in CSV/JSON with consistent keys.

Concrete improvement proposals (engineering roadmap)

  • Publish all KPI and budget tables as UTF-8 CSV and JSON, and register datasets in a central catalog (DCAT-AP compatible)
  • Add OpenAPI endpoints for frequently updated datasets, with pagination and CORS enabled for web apps
  • Provide schema and example payloads (JSON Schema) and maintain dataset versioning (semantic dates)
  • CI for data pipelines: unit tests to detect schema drift, encoding issues, missing daily rows
  • Developer portal + sample code (Python, JS) showing how to ingest data; offer CSV webhooks for streaming updates
  • Encourage standard date/time (ISO 8601), canonical column names, and machine-readable metadata fields (license, update frequency)
  • Implementation sketch: short tech stack

    • Storage: S3-compatible object store for raw dumps
    • Catalog: CKAN or a lightweight DCAT endpoint
    • API: FastAPI + OpenAPI autogenerated docs
    • ETL: Airflow + pandas for scheduled normalization; tests in pytest

    まとめ

    PDF-only publication is the main blocker for civic tech reuse. There are good signs — some ministries publish CSVs — but we need a unified practice: CSV/JSON + OpenAPI + metadata catalog. エンジニア的に言うと、API一本で解決する話なんですよね。これやれば政策の透明性と再現性がぐっと上がります!

    おかむーから一言

    I've built and shipped data products in GovTech; trusting APIs beats PDFs every time. Let's push for machine-readable policy data — tech can make democracy faster and clearer!