Code-Driven Manifesto: Making Japan's Administrative Data Truly Machine-Readable

IT Policy Proposals
Code-Driven Manifesto: Making Japan's Administrative Data Truly Machine-Readable

どうも〜おかむーです! Today I want to talk about something nerdy but civic: how Japan's government and local authorities publish data, and what we can do to make it actually useful for engineers and citizens.

  • The government's 2026 machine-readability rules push toward CSV/Excel, but most useful datasets still hide in PDFs or bespoke portals.
  • From an engineering view, the gap is API + schema + identifiers, not just "open" vs "closed".
  • Practical wins: canonical CSV/JSON schemas, versioned APIs, and simple CI checks will unlock reuse fast.

結論

Japan has clear policy momentum (see Digital Reform/内閣官房 and Digital Agency docs at digital.go.jp) and local system standardization work (soumu.go.jp), but implementation lags: many datasets remain in PDF or portal UIs. Engineer-wise, this is solvable with focused technical standards (CSV/JSON-first, stable APIs, example code) and tooling to automate quality checks.

Report

What I saw

Look at this: the Digital Reform conference decided machine-readability rules (PDF/summary at digital.go.jp). The rulebook even defines Level 1 as CSV/Excel readable by machines. Great! But in practice municipalities still publish procurement, budget, and performance results via HTML dashboards or PDFs (e.g. Miyagi procurement portal). That hurts reuse.

Why it matters (エンジニア的に言うと)

  • PDFs break pipelines: scraping is brittle and costly.
  • Portal-only UIs prevent bulk access and increase latency for civic apps.
  • No canonical schema = each project re-maps the same fields repeatedly.

Technical checklist (practical)

  • Data format: Provide CSV/JSON (JSON-LD or JSON API) as primary download; PDFs as secondary.
  • API: RESTful endpoints with pagination, filtering, and ETag/Last-Modified for caching.
  • Schema: Publish JSON Schema or CSV headers + data dictionary (column types, codes, identifiers).
  • Provenance: include publisher, timestamp, and versioning in metadata (schema.org/CSVW helps).

Quick code patterns

This is how a simple Python consumer could prefer CSV, fallback to PDF extraction:

import requests, pandas as pd

url_csv = 'https://example.go.jp/data.csv'

try:

df = pd.read_csv(url_csv)

except Exception:

# fallback: fetch PDF and use tabula or camelot

r = requests.get('https://example.go.jp/report.pdf')

with open('tmp.pdf','wb') as f: f.write(r.content)

# table = tabula.read_pdf('tmp.pdf', pages='all')

要するに〜: if data comes as CSV/JSON, the above consumer is one-liner; if PDF, you need heavy toolchain.

Policy vs Reality (numbers)

The Digital Agency timeline targets broader machine-readability by 2026 (see meeting materials). But current observables: many local gov systems are moving to standardized backends (gov cloud, soumu.go.jp initiatives) — progress exists but uneven. Measure progress by percent of datasets with CSV/JSON + API per prefecture.

Improvement roadmap

  • Mandate downloadable CSV/JSON for all datasets referenced in policy documents.
  • Provide reference API implementations (open-source) and CI validators that check schema and machine-readability before publication.
  • Offer grant support to small municipalities to migrate legacy systems to shared platforms (re-using gov cloud patterns).

まとめ

The policy framework is coming together — machine-readability rules and system standardization are in play. The remaining gap is largely technical: produce canonical CSV/JSON, stable APIs, and simple tooling for validation. Do that and civic tech can build real services fast.

おかむーから一言

I’ve built startups and shipped gov-facing systems: this is solvable if we standardize and share code. Tech + policy + a few good reference repos = huge public value. Let’s ship it!