Code-Driven Manifesto: Making Japan's Public Data Truly Machine-Readable

IT Policy Proposals
Code-Driven Manifesto: Making Japan's Public Data Truly Machine-Readable

どうも〜おかむーです! Today I'm taking a coder's lens to public-sector data in Japan — the place where PDFs go to retire and APIs get ghosted. エンジニア的に言うと、データが使えないと政策も検証できないんですよね。

  • Local and central gov are moving toward machine-readable rules, but many datasets are still PDFs.
  • The Digital Agency and Cabinet Secretariat have put concrete rules and decisions on machine readability — implementation gaps remain.
  • Technical fixes are straightforward: standard schema, a central API gateway, and sane metadata + tooling.

結論

Japan's policy direction is correct: require CSV/Excel/API-first publishing (see Digital Agency drafts at digital.go.jp and Cabinet Secretariat materials at cas.go.jp). However, the reality across municipalities is mixed—lots of PDF-locked reports and inconsistent schemas. 要するに、ルールはあるけど実装が追いついてないということです。エンジニア的には、API一本で解決する話なんですよ。

What I examined

Sources: Digital Agency's machine-readability rules draft (digital.go.jp), Cabinet Secretariat's digital governance pages (cas.go.jp), and GovTech Tokyo case studies showing practical reuse (govtechtokyo.or.jp / note.govtechtokyo.jp).

Machine-readability: PDF vs CSV

これ見てくださいよ:政府資料はまだPDFで公開されていることが多い。Digital Agency's draft explicitly ranks formats (Level 1: viewable/transcribable, Level 2: machine readable like CSV/Excel) — so PDF-only is considered insufficient. 要するに、PDFは人間には読めてもコードには読めないってことです。

技術的問題点:

  • Parsing PDFs is brittle (tables span pages, encodings, merged cells).
  • No standardized field names across municipalities ("population" vs "pop" vs "人数").
  • Missing machine-readable metadata (provenance, update frequency, license).

API and format availability

Many gov datasets lack a consistent REST API or OpenAPI description. Where APIs exist, schemas differ and authentication approaches vary. From an engineering perspective, that increases integration cost exponentially.

Quick code example: measure municipal compliance rate

import requests, pandas as pd

Hypothetical catalog endpoint returning JSON metadata

catalog = requests.get('https://example-gov.jp/api/datasets').json()

rows = []

for ds in catalog['datasets']:

fmt = ds.get('format','').lower()

rows.append({'id': ds['id'], 'format': fmt})

df = pd.DataFrame(rows)

rate = (df['format'].isin(['csv','xlsx'])).mean()

print(f"Machine-readable rate: {rate:.0%}")

要するに、まずはカタログからフォーマットを数えるだけで現状把握できます。

PDF rescue: pragmatic ETL patterns

  • Use table-extraction tools (tabula-py / camelot) as a stopgap.
  • Prefer automated OCR + heuristic column mapping, but log manual review.
  • Store canonical CSV in a versioned S3 bucket and publish an API layer on top.

Example transform pipeline:

  • Ingest: fetch PDF from lg.jp site
  • Extract: camelot -> raw CSV
  • Normalize: map column names to standard schema (use mapping table)
  • Validate: run schema checks (great_expectations)
  • Publish: push to API endpoint + metadata registry (DCAT-compliant)

Policy targets vs reality

The Cabinet Secretariat and Digital Agency have set rules and meeting outputs to improve machine-readability, but without enforcement timelines many municipalities lag. So numeric targets (percentage of datasets machine-readable) should be tracked via a central dashboard — GovTech Tokyo's dashboards are a good template.

Improvements I recommend (concrete)

  • Mandate an open catalog (DCAT) under a gov-central endpoint (go.jp). Include format, schema link, license, and update cadence.
  • Require OpenAPI/Swagger for any dataset API and publish examples.
  • Provide a reference Python/Node SDK and CI templates so municipalities can publish with one click.
  • Deploy a central ETL toolkit (containerized) that municipalities can reuse to convert PDFs to canonical CSV + tests.
  • Measure and publish a monthly "machine-readable rate" on a public dashboard (e-Stat style).

まとめ

Policy rules are landing — Digital Agency and Cabinet Secretariat materials show momentum — but the hard work is operational: standard schemas, central catalog, APIs, and reusable tooling. エンジニア的に言うと、ここは改善余地ありまくり!でも技術的対処法は明確で、やればかなり成果が出ます。

おかむーから一言

I've built products and startups that ship under messy constraints — this is solvable. Tech + policy aligned = measurable public value. Let's make government data actually usable, one API at a time!