Code-Driven Manifesto: Assessing Japan's Government Data Stack

IT Policy Proposals
Code-Driven Manifesto: Assessing Japan's Government Data Stack

Hey hey — okamu here! Today I’ll take a coder’s microscope to how Japanese national and municipal governments publish data. Short, punchy, and nerdy — just how I like it!

  • Governments offer growing catalogs (e-Gov, Tokyo Open Data, Digital Agency), but formats vary wildly
  • Machine-readability gaps: PDFs and ad-hoc CSVs still common, APIs exist but uneven coverage
  • Fixes: standardized APIs, OpenAPI docs, schema-first CSV publishing, and simple monitoring pipelines

結論

Japan has solid building blocks (e-Gov API catalog, Tokyo Open Data, Soumu guidance), but engineering practices lag: inconsistent formats, missing schema/metadata, and limited API coverage reduce reuse. 要するに、政策は立派でも“データプロダクト”になってないんですよね。API-first・schema-drivenで一気に改善できます!

Report

Background

This piece inspects public endpoints and datasets like e-Gov API (https://www.e-gov.go.jp/digital-government/api), Tokyo Open Data API (portal.data.metro.tokyo.lg.jp/opendata-api), Digital Agency case studies, Soumu open-data guidance, and sample CSV endpoints (notice.go.jp/docs/status_notice.csv, env.go.jp CSVs). These are real assets — good news — but the devil is in the engineering details.

What I looked at (engineer lens)

  • Format: JSON API vs CSV vs PDF. PDFs break automation; CSVs are fine if well-specified.
  • Metadata: presence of machine-readable schema (JSON Schema / CSVW / DCAT).
  • Discoverability: catalog APIs, OpenAPI specs, and pagination.
  • Update cadence & provenance: timestamps, versioning, license.

Findings (examples)

  • e-Gov provides an API catalog — great — but many endpoints lack OpenAPI/Swagger specs. That makes client generation or SDKs painful.
  • Tokyo’s Open Data API offers facility lists (e.g., PublicFacility) in machine form. Nice! But some datasets are CSV downloads with ambiguous column names and no schema.json.
  • Several gov CSVs are available (notice.go.jp, env.go.jp), yet some critical reports are only in PDF or embedded tables. This prevents automated monitoring or cross-dataset joins.

Technical implications

  • Data consumers must write brittle ETL: column guessing, ad-hoc date parsing, encoding fixes (Shift_JIS vs UTF-8), header normalization.
  • Lack of schema means every pipeline needs defensive coding (try/except, heuristics), increasing maintenance cost.

Concrete code example (how I'd fetch + validate a CSV)

# Minimal example: fetch CSV, normalize, validate with pandera

import requests, pandas as pd

from io import StringIO

import pandera as pa

r = requests.get('https://notice.go.jp/docs/status_notice.csv')

text = r.content.decode('utf-8')

df = pd.read_csv(StringIO(text))

schema = pa.DataFrameSchema({

'id': pa.Column(int),

'title': pa.Column(str),

'updated_at': pa.Column(pa.Check(lambda s: pd.to_datetime(s, errors='coerce').notnull()))

})

valid_df = schema.validate(df)

Policy vs reality

Governments often set numeric goals for transparency or dataset counts, but numbers alone are deceptive. Better metrics:

  • % machine-readable (JSON/CSV/API) vs PDF
  • % datasets with schema + license
  • mean freshness lag (hours/days)
  • API latency / error rates

Improvements (practical roadmap)

  • Adopt schema-first publishing: require JSON Schema/CSVW for every dataset.
  • Publish OpenAPI specs for each API endpoint, host on portal with client generation.
  • Standardize encodings (UTF-8), datetimes (ISO8601), and paging conventions.
  • Provide sample SDKs (Python/JS) and a low-barrier API console for civic developers.
  • Monitor quality: automated checks that flag PDFs, missing license, or broken encodings.
  • まとめ

    Japan’s public data infra is promising but still a mixed bag. With targeted engineering standards — OpenAPI, schema publishing, encoding rules — we can turn policy documents into reusable data products. エンジニア的に言うと、API一本とスキーマひとつで社会実装の壁がぐっと下がるんですよ。

    おかむーから一言

    I’ve built and shipped bad-ass gov tech and startups — making data product-quality is the low-hanging fruit for public good. Let’s ship APIs, not PDFs!