Code-Told Manifesto: How Japan’s Public Data Stacks Up (and How to Fix It)

IT Policy Proposals
Code-Told Manifesto: How Japan’s Public Data Stacks Up (and How to Fix It)

どうも〜おかむーです! Today I’m taking a bit of a nerdy stroll through Japan’s public data landscape — looking at APIs, CSVs, and the gaps between policy promises and engineering reality.

  • Governments publish lots of datasets, but formats and APIs are inconsistent
  • Machine-readability varies: Tokyo’s Open Data API is neat, many others are CSV/PDF dumps
  • Tech fixes are straightforward: API-first, schema standards, CI for data

結論

Japan has the right ingredients — e-Gov API catalogs, Tokyo’s Open Data API, and many CSV endpoints — but the ecosystem lacks consistent machine-readable standards and operational maturity. 要するに、政策はあるけど実装で損してるってことです。

Report

What I looked at

I checked the e-Gov API catalog (https://www.e-gov.go.jp/digital-government/api) and Tokyo’s Open Data API docs (https://portal.data.metro.tokyo.lg.jp/opendata-api/). I also noticed multiple government CSV endpoints (examples: https://www.soumu.go.jp/main_content/000323625.csv and other go.jp CSV links). These are real, public resources — nice!

これ見てくださいよ: Tokyo’s API exposes useful endpoints like GET /PublicFacility (wheelchair-accessible toilets, etc.), which is immediately usable by apps. On the flip side, plenty of datasets are only offered as CSV files or PDFs scattered across ministries — machine-readable but inconsistent.

Technical assessment

  • API presence: Positive in central hubs (e-Gov, Tokyo), but adoption is uneven across municipalities. Some cities have full APIs; others just drop CSVs or PDFs.
  • Data formats: CSV is common and OK, but schemas differ (column names, encodings). PDFs are a dead-end for automation.
  • Discoverability: e-Gov’s API catalog helps, but DCAT/metadata coverage is spotty.
  • Operational practices: Few examples of versioning, pagination, or clear rate limits in public APIs.

Engineer-wise, this means lots of brittle ETL jobs: parse CSVs with ad-hoc logic, re-map fields, handle encoding surprises. 要するに、毎回データクレンジングから始める羽目になる。

Concrete code snippets

Here’s a tiny Python example to fetch a CSV and normalize it with pandas:

import pandas as pd

url = 'https://www.soumu.go.jp/main_content/000323625.csv'

df = pd.read_csv(url, encoding='utf-8')

normalize column names

df.columns = [c.strip().lower().replace(' ', '_') for c in df.columns]

print(df.head())

And Node.js fetch to call a JSON API:

const res = await fetch('https://portal.data.metro.tokyo.lg.jp/api/PublicFacility');

const json = await res.json();

console.log(json.results.slice(0,3));

Policy vs reality: gaps

Policy aims (digital government, open data reuse) are clear in the e-Gov and Digital Agency narratives, but real-world indicators show friction:

  • Lack of standardized schemas means civic apps duplicate mapping work
  • PDF-only publications block reuse entirely
  • No consistent SLAs means developers can’t rely on uptime or stable endpoints

Improvement proposals (practical!)

  • API-first: require new datasets to publish a JSON API with schema and versioning
  • Standard schema layer: adopt DCAT + JSON Schema + JSON-LD context per dataset
  • CI/CD for data: validate CSVs against schema on publish; fail fast for broken rows
  • Central registry: extend e-Gov catalog with machine-readable metadata and health checks
  • Example SDKs: publish small client libs (Python/JS) to lower adoption barriers

まとめ

Japan’s public data infra is on the right track — there are good APIs and plenty of raw datasets. But to unlock large-scale civic tech, we need consistent formats, schema governance, and engineering practices (versioning, CI, SLAs). エンジニア的に言うと、これAPI一本で解決する話なんですよね。

おかむーから一言

I’ve built and shipped platforms that depend on reliable public data — standardizing APIs and automating quality checks is low-hanging fruit. Let’s make government data as dependable as production APIs!