Code-Driven Manifesto: Auditing Local Government Data with Engineering Eyes

- How messy are municipal datasets in Japan? Spoiler: mixed, but fixable!
- Machine-readability (CSV vs PDF), schema consistency, and APIs make or break reuse.
- I walk through concrete technical fixes, sample code, and policy-to-data checks you can use today.
結論
Public data is already an incredible asset, but too often it's published in formats or structures that make real accountability and reuse hard. As an engineer, I think the gap isn't lack of data—it's lack of predictable, machine-first publishing (CSV/JSON + schema + API + metadata). Fix that, and you get reproducible audits, civic apps, and policy feedback loops almost for free.
Report
Hey — Okamu here! Today I'm taking a tech-first look at how Japanese local governments publish data, based on catalogs like Tokyo's (https://catalog.data.metro.tokyo.lg.jp/dataset) and examples from Digital Agency and cities like Hakodate and Niigata. エンジニア的に言うと、these are the chokepoints that matter for civic reuse.
What I checked — quick tour
- Tokyo Open Data Catalog: lots of CSV entries, but varying schemas and metadata.
- Digital Agency sample CSVs (hospital lists, NICTER notices): good examples of machine files, but sometimes inconsistent encoding/headers.
- Hakodate / Niigata: guides and CSV inventories show awareness of CSV, yet manuals warn about numeric commas and encoding.
これ見てくださいよ:many datasets are CSVs—which is great—but CSV alone isn't enough. Without schema, charset declaration (UTF-8 vs Shift_JIS), license tags, and update cadence, automated pipelines choke.
Common technical problems (and why they break things)
- Inconsistent field names across municipalities (e.g. lat vs latitude vs 緯度): your join fails.
- Encodings and BOMs: some CSVs in Shift_JIS or with BOMs cause parsing errors in Python/pandas unless you handle them.
- Missing machine-readable metadata (no DCAT, no CSVW, no JSON Schema): you can't validate data automatically.
- PDF-first publication: tables trapped in PDFs need OCR/table-extraction; error-prone and brittle.
- No API or rate-limited portals: developers resort to scraping, which is fragile.
要するに、these issues prevent reproducible audits and make small civic apps cost-inefficient.
Quick code examples
Fetch a CSV from an open catalog, normalize encoding, and validate schema (short Python sketch):
import requests
import pandas as pd
from io import BytesIO
url = 'https://catalog.data.metro.tokyo.lg.jp/dataset/xxxx/resource/yyyy/download/sample.csv'
resp = requests.get(url)
try utf-8 then shift_jis
for enc in ('utf-8','shift_jis'):
try:
df = pd.read_csv(BytesIO(resp.content), encoding=enc)
break
except Exception:
df = None
simple schema check
expected_cols = {'id','name','latitude','longitude'}
missing = expected_cols - set(df.columns.str.lower())
print('Missing cols:', missing)
For production, add JSON Schema or pandera checks, CI that runs on publish, and automated alerts when schemas change.
Policy numbers vs data reality
Many policy documents (e.g. Cabinet Office CSV snippets for macro indicators) publish targets or summaries, but rarely publish KPI time series in a machine-friendly way. If a city claims “increase public facility accessibility by X%”, that should be a time series: facility inventory + access features (ramp, elevator) as CSV/JSON with timestamps. Without that, you can't programmatically check progress.
Practical checklist for municipalities
- Always publish CSV/JSON plus a machine-readable schema (CSV-W or JSON Schema).
- Use UTF-8 and declare charset in metadata.
- Publish license (Creative Commons) and update frequency.
- Provide a simple REST API (OpenAPI spec) or CKAN with resource endpoints.
- Geodata: provide GeoJSON or WGS84 lat/lon fields and validate coordinates.
Concrete improvements (engineering playbook)
- Adopt DCAT and CSV-W so catalogs are indexable and validators can run.
- Git-backed data portal: store datasets in a repo, run CI tests on pull requests, auto-deploy to catalog.
- Provide diff endpoints or changelogs so researchers can track KPI changes (use RFC 7233-style ETag/If-None-Match).
- Offer example SDKs/snippets (Python, JS) that show how to consume data safely.
- Convert legacy PDFs with a documented pipeline: OCR → table-extract (Camelot/Tabula) → human review → publish CSV.
Example: API design suggestion
- GET /api/v1/facilities?city=koto&format=geojson
- Schema: id, name, type, address, latitude, longitude, accessibility_features[], updated_at
- Provide OpenAPI and sample curl/pandas snippet in the portal docs
まとめ
Local governments are publishing useful datasets, and some (Tokyo, Digital Agency samples, city CSV lists) are doing a lot right. But to turn publication into real accountability and civic innovation, we need machine-first standards: consistent schemas, declared encodings, licenses, APIs, and CI-backed data QA. Implementing this is mostly engineering effort, not political miracle—so let's ship it!
おかむーから一言
I built startups and shipped civic systems — the tech stack for transparency exists. If cities standardize on schema + API + CI, we get immediate returns: reproducible audits, better apps, and real feedback loops between citizens and policymakers. Let's code the manifesto into reality!
Sources
- https://catalog.data.metro.tokyo.lg.jp/dataset
- https://www.harp.lg.jp/opendata/dataset/79.html
- https://opendata.pref.tokushima.lg.jp/dataset/5036.html
- https://catalog.data.metro.tokyo.lg.jp/dataset/?res_format=CSV&tags=%E8%87%AA%E6%B2%BB%E4%BD%93%E6%A8%99%E6%BA%96%E3%82%AA%E3%83%BC%E3%83%97%E3%83%B3%E3%83%87%E3%83%BC%E3%82%BF%E3%82%BB%E3%83%83%E3%83%88&groups=c029
- https://www.city.niigata.lg.jp/shisei/seisaku/it/open-data/index.files/csv_manual_v1.1.pdf
- https://www5.cao.go.jp/npc/
- https://www.intec.co.jp/column/smartcity-08.html
- https://ja.wikipedia.org/wiki/%E5%85%AC%E5%85%B1
- https://www.digital.go.jp/resources/data_case_study_private
- https://adtechmanagement.com/minnadepr-column/2025/11/02/koukyou-toha/
- https://www5.cao.go.jp/j-j/wp/wp-je24/csv/d1-1-4.csv
- https://notice.go.jp/docs/status_nicter.csv
- https://www.digital.go.jp/assets/contents/node/basic_page/field_ref_resources/0066e8a8-6734-44ab-a9a9-8e09ba9cb508/xxxxxx_hospital.csv
- https://www.jinji.go.jp/content/900024615.csv
- https://www5.cao.go.jp/j-j/wp/wp-je24/csv/d1-1-42.csv
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.