Code-driven Manifesto: Assessing Japan's Government Data & APIs from an Engineer's Lens

IT Policy Proposals
Code-driven Manifesto: Assessing Japan's Government Data & APIs from an Engineer's Lens

どうも〜おかむーです! Today I'm taking an engineer-first look at how Japanese national and municipal governments publish data and APIs —「コードで語るマニフェスト」ってやつです。エンジニア的に言うと、データ公開の品質は政策実行力そのものなんですよ〜

  • Governments provide valuable datasets but format fragmentation hurts reuse
  • e-Gov / GovTech Tokyo have APIs and portals, yet PDFs and inconsistent schemas remain common
  • Technical fixes (API-first, JSON Schema, CI/CD for datasets) can close the gap fast

結論

Public-sector data in Japan has strong foundations (e-Gov portals, Tokyo/Open Data initiatives, Soumu strategy), but delivery often falls back to PDF or ad-hoc CSVs. 要するに、データはあるけどエンジニアが使える形で出ていない、ということです。API standardization, machine-readable defaults, schema governanceが必要です。

Report: landscape and technical findings

What exists today (quick tour)

  • e-Gov / Administrative API catalog (api-catalog.e-gov.go.jp) exposes government APIs for e.g., corporate data and administrative records — good move toward machine accessibility.
  • Tokyo Open Data portal (portal.data.metro.tokyo.lg.jp) and GovTech Tokyo services show active dashboarding and reuse projects.
  • Soumu's Open Data Strategy outlines standard API and data model work (情報流通連携基盤共通API).

これ見てくださいよ:ポータルはあるけど、中身は混在してます。CSV/JSONがある一方で、重要な報告書や補助金一覧がPDFでしか公開されないことが多いんです。

Technical problems observed

  • PDF-first publications block automation. Extracting tables from PDFs requires tools (tabula/camelot) and is brittle.
  • Inconsistent schemas: different municipalities label the same field differently (e.g., "address" vs "addr" vs "所在地").
  • Lack of OpenAPI/JSON Schema: many APIs are undocumented or use ad-hoc query parameters.
  • No dataset CI: changes to published CSVs break downstream pipelines unexpectedly.

Example: fetching corporate data (conceptual)

import requests

r = requests.get('https://api-catalog.e-gov.go.jp/endpoint/corporation?corporateNumber=1234567890123')

print(r.status_code, r.headers.get('content-type'))

expect application/json, then validate with JSON Schema

要するに、API一本で解決する話なんですよね。

PDF -> CSV technical patterns

  • Use tabula-py or camelot to extract tables
  • Apply post-processing: normalize encodings, date formats, kanji variants
  • Better: avoid the pain and publish CSV/JSON/GeoJSON alongside human PDFs

Policy targets vs achievements

Soumu's strategy and e-Gov APIs set an ambition for standard APIs and cross-domain data flows. But reality: many local government datasets remain non-machine-readable. That gap indicates implementation friction — limited tooling, legacy procurement, and operational capacity.

Recommendations (engineering roadmap)

  • API-first publishing: every dataset gets a JSON/CSV + OpenAPI/JSON Schema.
  • Central schema registry: canonical vocabularies (法人番号, address, fiscal-year) shared across LGs.
  • Dataset CI/CD: Git-backed pipelines (GitHub/GitLab) with tests, validators, and changelogs.
  • PDF as human artifacts only: attach machine-readable siblings as default.
  • Developer experience: interactive API consoles, rate-limited sandbox keys, and examples.
  • Monitoring and SLAs: dataset availability metrics (uptime, freshness) on dashboards.
  • まとめ

    現状はスタート地点に立っているものの、実運用で使いやすいデータ提供にはまだ道のりがあります。エンジニア視点での小さな改善(schema, API docs, dataset CI)が政策の実効性を大きく高めるんです。

    おかむーから一言

    起業家として、エンジニアとして言うと、政府データはプロダクトです。小さな改善を積み上げていけば、社会のインフラはもっと速く良くなりますよ〜