Code-Driven Manifesto: Evaluating Japan’s Public Data and APIs from an Engineer’s Lens

どうも〜おかむーです! Today I’m digging into how Japanese government and local authorities publish data — from e-Gov APIs to CSV drops on ministry sites — and what engineers actually need to make policy measurable and reusable.
- The ecosystem has real building blocks (e-Gov APIs, Tokyo Open Data portal, CSV dumps), but machine-readability and schema quality are inconsistent.
- Common pain: mixed encodings, PDF-embedded tables, and missing time-series KPIs that make policy verification hard.
- Fixable with API-first publishing, DCAT/JSON-LD metadata, standardized encodings, and simple ETL patterns.
結論
Japan has the right ingredients for a data-driven public sector, but delivery is uneven. エンジニア的に言うと、API一本ときちんとした schema があれば多くの政策議論は定量的に検証可能なんですよね。今あるCSVやAPIをつなげて、オープンで検証可能な“コードで語るマニフェスト”を作るのが現実的です。
Report: what's out there and what's missing
What I looked at (examples from search results)
- e-Gov Administrative API catalog (https://www.e-gov.go.jp/digital-government/api) — shows central gov moving toward APIs.
- Tokyo Open Data API (https://portal.data.metro.tokyo.lg.jp/opendata-api/) — solid catalog with endpoints like
PublicFacility. - notice.go.jp CSV endpoint (
/docs/status_notice.csv) and various*.csvfiles linked on ministry domains — evidence of direct CSV publishing. - Digital Agency case studies on private reuse — shows appetite for building on gov datasets.
これ見てくださいよ:there are APIs and CSVs, but formats vary, and metadata is often incomplete. That makes automated validation and cross-jurisdiction comparisons painful.
Technical pain points (concrete)
- Encoding and formats: many legacy CSVs may be Shift_JIS or not declare encoding; must detect and normalize to UTF-8.
- PDFs vs CSV: important tables are sometimes only in PDF (hard to extract reliably). 要するに、機械可読性が低いということです。
- Missing schema/metadata: datasets lack DCAT/Schema.org/JSON-LD descriptions, making discovery and type-checking brittle.
- Incomplete time-series/KPI publishing: policy targets exist, but machine-readable actuals (monthly/quarterly numbers) are often missing.
- Authentication and rate limits: e-Gov APIs exist but adoption increases require clear SLAs, example SDKs and client libraries.
Example: pragmatic ETL to turn a government CSV into usable JSON
import requests, chardet, pandas as pd
r = requests.get('https://notice.go.jp/docs/status_notice.csv')
enc = chardet.detect(r.content)['encoding']
text = r.content.decode(enc or 'utf-8', errors='replace')
from io import StringIO
df = pd.read_csv(StringIO(text))
Basic cleaning
df.columns = [c.strip() for c in df.columns]
Export normalized JSON-LD chunk
print(df.head().to_json(orient='records', force_ascii=False))
Notes: use chardet to detect encoding, validate columns, and then publish JSON-LD with schema.org types for interoperability.
Policy measurement: how to check targets vs reality
- Ask: does the policy publish a numeric target and a machine-readable time series? If not, you can’t programmatically verify.
- Practical approach: create a minimal data contract — e.g.,
policy_id,target_value,target_date,observed_value,observed_date,source_id. - Link datasets via stable identifiers (e.g., agency code + dataset ID). Then build dashboards that auto-refresh from APIs.
Governance & engineering improvements (concrete proposals)
- API-first publishing: require all new datasets expose a paginated JSON API and an OpenAPI spec. Examples: wrap existing CSV endpoints with a thin API layer.
- Metadata standardization: adopt DCAT-AP-JP + JSON-LD so portals like Tokyo & e-Gov expose discoverable machine metadata.
- Encoding policy: mandate UTF-8 for all public data, or include correct HTTP
Content-Type; charset=headers. - Provide sample SDKs and Postman collections for major APIs to lower adoption friction.
- Publish KPIs as machine-readable time series (CSV/JSON) and link them to the underlying datasets.
- CI for data quality: set up automated validators (schema checks, null-rate alerts, freshness tests) and show data health badges on catalogs.
まとめ
- Japan’s public data landscape is promising: e-Gov APIs and municipal portals exist, and private reuse cases show value.
- Main barriers are machine-readability, inconsistent metadata, and lack of KPI time-series for policy verification.
- Engineering fixes are straightforward: API-first, UTF-8, DCAT/JSON-LD metadata, simple ETL patterns, and automated data quality pipelines.
おかむーから一言
テクノロジーで行政はもっとオープンで検証可能になります!エンジニア目線の小さな改善を積めば、政策は数字で語れるようになるんですよ。やりましょう!
Sources
- https://www.jichi.ac.jp/
- https://www.e-gov.go.jp/digital-government/api
- https://www.jichi.ac.jp/web_text/
- https://portal.data.metro.tokyo.lg.jp/opendata-api/
- https://www.jichi.ac.jp/library/
- https://notice.go.jp/docs/status_notice.csv
- https://www.jinji.go.jp/content/900024615.csv
- https://www.env.go.jp/content/900398071.csv
- https://www.inpit.go.jp/content/100869372.csv
- https://www.mhlw.go.jp/content/001429362.csv
- https://kotobank.jp/word/%E5%85%AC%E5%85%B1-494676
- https://www.intec.co.jp/column/smartcity-08.html
- https://www.pref.miyagi.jp/soshiki/jyoho/miyapo.html
- https://www.digital.go.jp/resources/data_case_study_private
- https://miyagi.efftis.jp/04000/PPI/Public/public/common/jsp/OTeaP_main_frame.jsp
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.