Code-Driven Manifesto: Evaluating Japan’s Public Data and APIs from an Engineer’s Lens

IT Policy Proposals
Code-Driven Manifesto: Evaluating Japan’s Public Data and APIs from an Engineer’s Lens

どうも〜おかむーです! Today I’m digging into how Japanese government and local authorities publish data — from e-Gov APIs to CSV drops on ministry sites — and what engineers actually need to make policy measurable and reusable.

  • The ecosystem has real building blocks (e-Gov APIs, Tokyo Open Data portal, CSV dumps), but machine-readability and schema quality are inconsistent.
  • Common pain: mixed encodings, PDF-embedded tables, and missing time-series KPIs that make policy verification hard.
  • Fixable with API-first publishing, DCAT/JSON-LD metadata, standardized encodings, and simple ETL patterns.

結論

Japan has the right ingredients for a data-driven public sector, but delivery is uneven. エンジニア的に言うと、API一本ときちんとした schema があれば多くの政策議論は定量的に検証可能なんですよね。今あるCSVやAPIをつなげて、オープンで検証可能な“コードで語るマニフェスト”を作るのが現実的です。

Report: what's out there and what's missing

What I looked at (examples from search results)

  • e-Gov Administrative API catalog (https://www.e-gov.go.jp/digital-government/api) — shows central gov moving toward APIs.
  • Tokyo Open Data API (https://portal.data.metro.tokyo.lg.jp/opendata-api/) — solid catalog with endpoints like PublicFacility.
  • notice.go.jp CSV endpoint (/docs/status_notice.csv) and various *.csv files linked on ministry domains — evidence of direct CSV publishing.
  • Digital Agency case studies on private reuse — shows appetite for building on gov datasets.

これ見てくださいよ:there are APIs and CSVs, but formats vary, and metadata is often incomplete. That makes automated validation and cross-jurisdiction comparisons painful.

Technical pain points (concrete)

  • Encoding and formats: many legacy CSVs may be Shift_JIS or not declare encoding; must detect and normalize to UTF-8.
  • PDFs vs CSV: important tables are sometimes only in PDF (hard to extract reliably). 要するに、機械可読性が低いということです。
  • Missing schema/metadata: datasets lack DCAT/Schema.org/JSON-LD descriptions, making discovery and type-checking brittle.
  • Incomplete time-series/KPI publishing: policy targets exist, but machine-readable actuals (monthly/quarterly numbers) are often missing.
  • Authentication and rate limits: e-Gov APIs exist but adoption increases require clear SLAs, example SDKs and client libraries.

Example: pragmatic ETL to turn a government CSV into usable JSON

import requests, chardet, pandas as pd

r = requests.get('https://notice.go.jp/docs/status_notice.csv')

enc = chardet.detect(r.content)['encoding']

text = r.content.decode(enc or 'utf-8', errors='replace')

from io import StringIO

df = pd.read_csv(StringIO(text))

Basic cleaning

df.columns = [c.strip() for c in df.columns]

Export normalized JSON-LD chunk

print(df.head().to_json(orient='records', force_ascii=False))

Notes: use chardet to detect encoding, validate columns, and then publish JSON-LD with schema.org types for interoperability.

Policy measurement: how to check targets vs reality

  • Ask: does the policy publish a numeric target and a machine-readable time series? If not, you can’t programmatically verify.
  • Practical approach: create a minimal data contract — e.g., policy_id, target_value, target_date, observed_value, observed_date, source_id.
  • Link datasets via stable identifiers (e.g., agency code + dataset ID). Then build dashboards that auto-refresh from APIs.

Governance & engineering improvements (concrete proposals)

  • API-first publishing: require all new datasets expose a paginated JSON API and an OpenAPI spec. Examples: wrap existing CSV endpoints with a thin API layer.
  • Metadata standardization: adopt DCAT-AP-JP + JSON-LD so portals like Tokyo & e-Gov expose discoverable machine metadata.
  • Encoding policy: mandate UTF-8 for all public data, or include correct HTTP Content-Type; charset= headers.
  • Provide sample SDKs and Postman collections for major APIs to lower adoption friction.
  • Publish KPIs as machine-readable time series (CSV/JSON) and link them to the underlying datasets.
  • CI for data quality: set up automated validators (schema checks, null-rate alerts, freshness tests) and show data health badges on catalogs.

まとめ

  • Japan’s public data landscape is promising: e-Gov APIs and municipal portals exist, and private reuse cases show value.
  • Main barriers are machine-readability, inconsistent metadata, and lack of KPI time-series for policy verification.
  • Engineering fixes are straightforward: API-first, UTF-8, DCAT/JSON-LD metadata, simple ETL patterns, and automated data quality pipelines.

おかむーから一言

テクノロジーで行政はもっとオープンで検証可能になります!エンジニア目線の小さな改善を積めば、政策は数字で語れるようになるんですよ。やりましょう!