Code-Driven Manifesto: Evaluating Japan's Local Gov Data Through an Engineering Lens

IT Policy Proposals
Code-Driven Manifesto: Evaluating Japan's Local Gov Data Through an Engineering Lens

どうも〜おかむーです! Hi — today we're doing a quick, nerdy check on how Japanese local governments publish data and what that means for policy promises. エンジニア的に言うと、データ公開は政策実行のAPIなんですよ〜

  • Local govs publish lots of datasets, but formats and metadata are inconsistent
  • PDF-locked statistics and missing APIs hamper reuse and measurable policy follow-through
  • Practical fixes: standardize CSV/metadata (DCAT), offer simple APIs, and publish provenance

結論

Open data policy momentum exists (see Tokyo Catalog, Niigata CSV guidance, Hakodate CSV list), but machine-readability and API-first practices are still spotty. 要するに、政策の「数値目標」は出るけど、検証できる形で出てこないことが多いんです!

Report: what I checked and what it means

What I looked at

  • Tokyo Open Data Catalog (catalog.data.metro.tokyo.lg.jp) — a real catalog, many CSV/XLSX files
  • Niigata CSV manual (city.niigata.lg.jp) — shows local efforts to standardize CSVs
  • Hakodate / Saitama portals — examples of CSV lists and dataset catalogs
  • GovTech Tokyo materials on data utilization and dashboards

These are good signals: gov domains (go.jp / lg.jp) and portals exist. But look at this: many municipalities still publish key reports as PDFs or mixed Excel sheets. This kills automation.

Technical problems I see

  • PDF vs CSV: PDFs embed tables that need OCR/tabula to extract. That adds error and manual steps.
  • Metadata absence: no DCAT/JSON-LD to describe schema, license, update cadence.
  • No stable APIs: catalog pages list files but rarely provide documented REST APIs or OpenAPI specs.
  • Inconsistent CSV schemas: different column names, encodings (Shift_JIS vs UTF-8), date formats.

These lead to reproducibility problems: you can't easily write a pipeline that pulls monthly metrics across cities.

Concrete code examples (how I'd approach it)

  • Quick normalization (Python/pandas):
import pandas as pd

url = 'https://example.city/data/population.csv'

df = pd.read_csv(url, encoding='utf-8')

normalize columns

df = df.rename(columns=lambda c: c.strip().lower().replace(' ', '_'))

df['date'] = pd.to_datetime(df['date'])

df.to_parquet('population.parquet')

  • Extracting from PDF (tabula-py):
from tabula import read_pdf

tables = read_pdf('report.pdf', pages='all', lattice=True)

manual QA required

  • Use national API where possible (e-Stat) via API key and aggregate with local CSVs for comparability.

Policy vs Implementation gap

Many local promises say "open data will improve transparency" but the output is often human-readable only. Without machine-readable timestamps, primary keys, and licensing, you can't measure progress programmatically. 要するに、政策評価がスッとできないんです。

Practical roadmap (engineering-first)

  • Mandate CSV/JSON-first publishing, UTF-8, RFC4180-compatible
  • Require dataset metadata: DCAT or JSON-LD with license, update frequency, schema
  • Provide lightweight REST endpoints or a CKAN-backed catalog with CORS enabled
  • Publish sample OpenAPI spec and 1-line curl examples for each dataset
  • Offer migration tools: server-side converters from XLS/PDF to canonical CSV and QA pipelines

Example minimal API response pattern (JSON):

{

"dataset_id": "population",

"title": "Population by age",

"updated": "2025-01-01",

"download_url": "https://.../population.csv",

"schema": {"columns": ["date","age_group","count"]}

}

まとめ

Local governments in Japan are on the right track—catalogs and CSV guidance exist—but the devil is in the machine-readability. Fixing encoding, metadata, and offering simple APIs turns policy text into testable, auditable code. This is how manifestos become measurable.

おかむーから一言

I love that Japan has momentum on GovTech—let's make sure the data isn't just pretty for humans but friendly for code too. Tech can make policy accountable, and that's what I'm pushing for!