Code-driven Manifesto: Diagnosing Japan's Open Data, PDFs and APIs

IT Policy Proposals
Code-driven Manifesto: Diagnosing Japan's Open Data, PDFs and APIs

どうも〜おかむーです! Today I'm going to take a practical, engineering-first look at how central and local Japanese government data is published — and what we can do with it. エンジニア的に言うと、このレポートは「_policy as data_」をコードで検証する試みです。

  • Governments publish lots of useful figures, but formats and access patterns vary wildly
  • Machine-readable availability (CSV / API / JSON-LD) is mixed; many PDFs remain the source of truth
  • Small fixes (CSV, schema, API, provenance) unlock big reuse — here's how

結論

Governments already publish many datasets (see digital.go.jp, cas.go.jp, mhlw.go.jp, env.go.jp), but inconsistent formats (PDF vs CSV), lack of stable APIs, and weak metadata hamper reuse. 要するに、機械可読性とAPI中心の公開戦略を優先すれば、政策評価も民主的な監視もずっとスムーズになります。

Report

What I looked at

これ見てくださいよ — Cabinet Secretariat guidance on "官民におけるデータの利活用" (https://www.cas.go.jp) and Digital Agency notes on "Data for AI" (digital-gov.note.jp / digital.go.jp) discuss AI-ready, machine-readable admin data. At the same time, you can find many CSV endpoints on various .go.jp domains (examples surfaced in search results: notice.go.jp CSV, mhlw.go.jp CSV, env.go.jp CSV files).

PDF vs CSV vs API: the practical differences

  • PDF: great for human-readable reports (KPI summaries, narrative). Bad for automation.
  • CSV: much better — tabular, easy to ingest with pandas or R.
  • API/JSON: best for real-time, versioned, documented access.

エンジニア的に言うと、PDFは"serialised human UI"であって、APIは"behavioural contract"なんですよね。

Concrete spot-check: KPI publishing for grant programs

The Digital Grand (デジタル田園都市国家構想交付金) documentation (prefecture PDFs and central guidelines) sets KPIs and reports outcomes. But often the KPI tables are embedded in PDFs (see chiso.go.jp guideline PDF and local PDF results). That makes automated evaluation — e.g., yearly progress vs target — a manual pain.

Quick code example — turning a CSV into KPI checks

import pandas as pd

example: fetch a published CSV (one of the gov CSV URLs)

url = 'https://www.env.go.jp/content/900398071.csv'

df = pd.read_csv(url, encoding='utf-8')

assume df has columns: year, project_id, kpi_name, target, actual

summary = df.groupby('kpi_name').agg({'target':'sum','actual':'sum'})

summary['achievement_rate'] = summary['actual'] / summary['target']

print(summary.sort_values('achievement_rate'))

要するに、CSVならこのコードで数分で評価できるんです。

Data quality & metadata

Observed issues:

  • Missing schema: no machine-readable schema or datatypes.
  • Encoding & normalization: mixed encodings and header formats across agencies.
  • Provenance: which version of the KPI does this CSV reflect? No versioning.

These are solvable with simple conventions: JSON Schema / CSVW, UTF-8 normalization, and dataset-level metadata (modified date, source URL, license).

API design suggestions

  • Provide RESTful endpoints per dataset with OpenAPI spec.
  • Support filtered queries (year, region, KPI) and pagination.
  • Include machine-readable metadata (schema + provenance) in dataset header.
  • Offer CSV/JSON/JSON-LD outputs and an S3-like stable artifact store for snapshots.

Example small roadmap

  • Identify high-value PDF tables (grants, KPI results) and prioritize CSV extraction.
  • Publish CSV + JSON Schema + OpenAPI for each dataset.
  • Automate extraction from legacy PDFs using extraction pipeline (tabula/pdftotext + human review) and publish snapshots.
  • Add continuous integration: every change triggers schema checks and sample queries.
  • まとめ

    政府はデータを出しているけど、フォーマットとアクセスの揺れで利活用が遅れている。APIと機械可読性(CSV/JSON Schema/metadata)を優先すれば、政策評価・AI利用・市民監視が一挙に改善します。小さな投入で大きく効く領域です。

    おかむーから一言

    データをちゃんと公開すれば、政策はもっと速く良くなる。テクノロジーで社会をアップデートするって、結局こういう地道な積み重ねなんですよね!