Code-Speaks Manifesto: Auditing Japan’s Local Government Data Stack

IT Policy Proposals
Code-Speaks Manifesto: Auditing Japan’s Local Government Data Stack

どうも〜おかむーです! Hey — Okamu here! Today I take a code-first look at how Japanese national and local governments publish data, what’s actually machine-readable, and how we could turn policy promises into developer-friendly infrastructure.

  • This piece inspects official programs (総務省, デジタル庁) and their open-data outputs
  • Main gap: lots of PDF and inconsistent schemas; few uniform APIs across municipalities
  • Practical fixes: publish CSV/JSON + schema, OpenAPI endpoints, CI for data quality, and a federated catalog

結論

Japan has the right policy building blocks — see 総務省’s自治体情報システムの標準化・共通化 and デジタル庁’s Data for AI initiative — but the implementation surface is messy: many datasets are PDFs, formats vary wildly, and APIs are fragmented. 要するに、policy exists, but engineers still waste time PDF-scraping instead of building apps. We can fix this with targeted engineering: machine-readable formats, schemas, OpenAPI, and CI pipelines for datasets!

Report: what I looked at and what I found

What I used as anchors:

  • 総務省: 自治体情報システムの標準化・共通化 (https://www.soumu.go.jp/...) — shows the standardization push and mentions a PMO progress-tracking tool
  • デジタル庁: local data case studies and Data for AI (https://www.digital.go.jp/, https://digital-gov.note.jp/...) — signals an explicit push to make administrative data AI-ready
  • 内閣官房資料 on data reuse & machine readability (https://www.cas.go.jp/...) — highlights machine-readable rules

These official pages are good policy references, but when you crawl municipal portals you frequently hit three problems:

1) PDF-first publication: datasets or tables embedded in PDFs rather than CSV/JSON

2) No consistent metadata: missing schema, no DCAT/JSON-LD, license fields inconsistent

3) Sparse APIs: some metropolitan cities have APIs, but most towns simply host files or PDFs

This is consistent with the push described on the Digital Agency and MIC pages: the intent is there, but the execution across ~1,700 municipalities is uneven.

Technical audit: formats, access patterns, and developer ergonomics

This is an engineer's checklist of what matters:

  • File types: CSV / JSON / GeoJSON = great. XLSX = okay (but treat as binary with clear schema). PDF = pain.
  • Metadata: catalog-level metadata (DCAT), dataset-level schema, license (CC-BY or similar), update cadence field.
  • API: OpenAPI spec, CORS enabled, rate limits documented, stable base path and versioning.
  • Provenance: dataset creation timestamp + last update + unique dataset id.

This is what I see in the wild: many local sites provide PDFs for council budgets, facility lists, or disaster reports. Some prefectures have decent CSVs or small REST endpoints, but they’re not standardized.

Code example: quickly reading a municipal CSV or falling back to PDF scraping (Python pseudocode)

import requests

import pandas as pd

from tabula import read_pdf # or camelot

url_csv = 'https://example-city.gov/data/facilities.csv'

resp = requests.get(url_csv)

if resp.status_code == 200 and 'text/csv' in resp.headers.get('content-type',''):

df = pd.read_csv(url_csv)

else:

# fallback: extract table from PDF

pdf_url = 'https://example-city.gov/data/facilities.pdf'

tables = read_pdf(pdf_url, pages='1-3')

df = pd.concat(tables)

print(df.head())

Tabula/Camelot pipelines work, but they’re brittle. You want canonical CSV/JSON with a schema so you can skip brittle OCR steps.

Measuring the gap: policy targets vs reality

Digital policy documents (e.g. Data for AI) emphasize machine-readable administrative data. To operationalize that, we should track simple KPIs per municipality:

  • % of datasets provided in CSV/JSON/GeoJSON
  • # of datasets with machine-readable schema (JSON Schema / CSVW / DCAT)
  • # of datasets with explicit license and update cadence
  • # of OpenAPI endpoints

A short script can crawl city open-data pages and categorize file types and metadata fields — this gives a repeatable compliance metric.

Practical improvements — concrete, developer-friendly steps

1) Mandate dataset-level machine-readable metadata

- Use DCAT/JSON-LD and include: title, description, license, publisher, updateFrequency, schema

2) Promote CSV/JSON as first-class artifacts

- If PDF is the legal artifact, still publish CSV/JSON proxies and keep PDFs as human-readable proofs

3) Require an OpenAPI gateway for standard endpoints

- Minimal spec: /datasets, /datasets/{id}/download, /search?q=, /health

4) Implement CI for datasets

- Example: use Great Expectations to assert column types, value ranges, no nulls in keys

5) Provide code snippets and sample notebooks per dataset

- One notebook per dataset = huge UX win for reuse

6) Centralized federated catalog

- Each municipality can publish a DCAT endpoint; a central aggregator regularly harvests and indexes

Example: Great Expectations-style assertion (pseudocode)

expectations:

- expect_column_values_to_not_be_null: ["facility_id"]

- expect_column_values_to_be_between:

column: capacity

min_value: 0

System design sketch: federated catalog + gateway

  • Each municipality exposes a small catalog endpoint (/dcat.json) and dataset files (CSV/JSON)
  • Central harvester polls catalogs, validates schemas, and publishes a unified index
  • API gateway provides unified OpenAPI, search, and per-dataset download links
  • CI checks run on harvest to flag schema drift

This architecture follows patterns used in other countries and fits the Japanese legal/operational landscape: local autonomy + central coordination.

Use cases unlocked

  • Real-time disaster dashboards combining elevation, river gauge, and facility shelter lists
  • Flood insurance or mortgage risk scoring using local building stock + hazard maps
  • Small businesses building local mobility apps using bus timetables in GTFS (if municipalities publish it!)

The example in the Domestic Smart City article (sorabatake) shows how combining elevation and open damage data produced a public-facing flood risk map — imagine if municipal data were standard across prefectures!

まとめ

  • Policy momentum is real (総務省・デジタル庁), but implementation is fragmented and PDF-heavy
  • Engineers need: CSV/JSON + schema, DCAT catalogs, OpenAPI endpoints, CI data tests, and a federated catalog
  • Quick wins: require machine-readable metadata, publish sample notebooks, and run automated validators on harvest

If you’re building services on top of municipal data, invest in a small ingestion stack (harvester + validator + snapshots). If you’re a policymaker, set simple, measurable KPIs like % machine-readable datasets and require DCAT endpoints.

おかむーから一言

I’ve built products that die on bad inputs — data quality is code quality. Let’s make government data behave like good APIs: versioned, documented, and testable. Tech can make policy measurable, and that’s where the magic happens!