Code-driven Manifesto: Evaluating Japan’s Gov Data Practices from an Engineer’s Lens

IT Policy Proposals
Code-driven Manifesto: Evaluating Japan’s Gov Data Practices from an Engineer’s Lens

どうも〜おかむーです! Today I want to do a bit of a nerdy inspection of how government and municipal data are actually published and used in Japan — with an "engineer-first" lens. Code speaks, policies don’t just sound good on paper, they should be verifiable by data and machines!

  • This piece inspects machine-readability, APIs, and KPI reporting for national/local projects (e.g. Digital田園都市, GovTech Tokyo).
  • Key problems: data stuck in PDFs, inconsistent formats, sparse APIs; engineering fixes can deliver transparency fast.
  • I propose concrete tooling and a tiny code example to turn PDFs into usable CSV/JSON and recommend API design + CI for data quality.

結論

Government has the right priorities — dashboards and KPI frameworks are showing up (see Digital庁, GovTech東京) — but the implementation is half-done: many KPIs and project results are locked in PDFs, inconsistent spreadsheets, or scattered sites (chisou.go.jp, digital.go.jp, local city pages). 要するに、政策が数値で語れる状態にするためには「機械可読性」と「継続的データパイプライン」が必要です。

Report: what I checked and what I found

Why I looked: as a GovTech practitioner I want policies to be auditable by developers and analysts. I scanned public entry points: Digital庁 (digital.go.jp), Cabinet Office guidance (chisou.go.jp), GovTech Tokyo (govtechtokyo.or.jp), and a local example (City of Sukagawa: city.sukagawa.fukushima.jp) referenced in public evaluation pages.

What’s common

  • PDF-first reporting: evaluation reports and KPI tables are often embedded in PDFs (check chisou.go.jp guideline documents). That makes automated verification hard.
  • Mixed formats: CSV sometimes exists but often inconsistent schemas, messy headers, and date formats.
  • Sparse APIs: a few dashboards exist (GovTech Tokyo dashboards) but many projects lack stable, documented REST APIs or data catalogs.

Technical implications (engineer-y)

  • PDF → data friction: OCR and table extraction aren't reliable for ongoing monitoring. エンジニア的に言うと、PDF is a snapshot format, not a streaming data source.
  • No canonical ID: entities (projects, grants) lack persistent identifiers across datasets, so joins are painful.
  • Missing metadata: no machine-readable schemas (JSON Schema, Data Packages), so consumers must guess column meanings.

Concrete example: Digital田園都市 KPIs

The program (see chisou.go.jp guidance on Check/Action) defines KPIs and monitoring but published evaluations are PDF reports. That means:

  • It's hard to track trend programmatically across fiscal years.
  • Auditors and citizens can’t easily reproduce numbers in dashboards without manual scraping.

This is not an accusation — it’s a technical debt story.

Practical fixes (short-term → long-term)

Short-term

  • Publish CSV/JSON next to every PDF. Even a single consistent CSV per KPI per fiscal year unlocks reproducibility.
  • Provide simple REST endpoints (e.g. /api/v1/kpis?fiscal=2024®ion=sukagawa) with pagination and schema discovery.
  • Use a light catalog (e.g. frictionless Data Package) so consumers can read schema and provenance.

Medium-term

  • CI pipelines for data: treat data like code. Use tests (row counts, schema, ranges) on every data update.
  • Persistent IDs: assign project/grant IDs and expose them across datasets (grants, outputs, audits).
  • Open API spec: publish OpenAPI/Swagger and example queries.

Long-term

  • Standardize on JSON-LD or Schema.org for open data, register datasets in a central portal (e-Stat integration), and run automated KPI checks.

Code example — extracting a table from a PDF and publishing it as CSV/JSON

This is a tiny Python sketch that engineers can run as a first-aid to convert PDFs into analyzable CSVs. Code people will get this — it's not production but a reproducible start.

import tabula

import pandas as pd

Extract tables from a PDF report (first page as example)

requires: pip install tabula-py (and Java), or use Camelot

tables = tabula.read_pdf("evaluation_report.pdf", pages=1, multiple_tables=True)

Normalize and export

df = tables[0]

quick clean: unify column names

df.columns = [c.strip().lower().replace('\n','_') for c in df.columns]

export

df.to_csv("kpi_extract_2024.csv", index=False)

print('Exported CSV with', len(df), 'rows')

Better: wrap this in a GitHub Action that runs on new PDF releases, validates row counts and value ranges, and commits CSVs to a data repository.

API design suggestion (minimal)

  • GET /api/v1/kpis — returns list of KPI definitions (JSON Schema)
  • GET /api/v1/kpis/{kpi_id}/observations?from=2020-04-01&to=2025-03-31®ion=XXX — returns time-series
  • GET /api/v1/projects/{id} — returns project metadata (with persistent id, budget, outputs)

Security and privacy: public KPIs should be fully open; personal data must be anonymized before release. Use rate limits but prefer open access for transparency.

Open data tooling to adopt

  • Frictionless Data (data package + goodtables)
  • CKAN or a lightweight static-data catalog for hosting
  • GitHub Actions + Great Expectations for CI data tests
  • Postgres/Timescale or data lake for storing time-series of KPIs

まとめ

This is doable and cheap: start by publishing CSVs alongside PDFs and adding an opinionated API. 要するに、政策の信頼は「人が読めること」と「機械が検証できること」の両方で成り立っているんです。GovTech Tokyo and Digital庁 are building the right muscles; now it’s time to operationalize them with data pipelines, IDs, and automated checks.

おかむーから一言

I’ve built startups and shipped GovTech products — tech can make transparency non-negotiable. Start with CSV+API+CI, and you’ll turn PDF drama into reproducible policy truth faster than you think!