Code-Driven Manifesto: How Japan’s Local Open Data Stacks Up (and How to Fix It)

IT Policy Proposals
Code-Driven Manifesto: How Japan’s Local Open Data Stacks Up (and How to Fix It)

どうも〜おかむーです! Today I want to dig into how Japanese local governments publish data — the good, the messy, and the fixable — from an engineer's point of view. Let's go!

  • Many municipalities publish CSVs (Tokyo catalog, Saitama portal, Hakodate lists) but formats vary wildly.
  • Machine-readability wins when data is CSV+schema; PDFs and ad-hoc columns break reuse.
  • Practical fixes: API endpoints, schema validation, DCAT metadata, CI for data — small engineering moves with big civic ROI.

結論

Local governments are publishing valuable datasets (see Tokyo, Saitama, Niigata guidance) but lack consistent machine-readable schemas and APIs. 要するに、データがあるのに使いにくい。エンジニアの観点からは、CSV+schema+APIという三点セットを整備すれば、利活用が一気に進みます!

Report: what I checked and what it means

What the catalogs show

  • Tokyo Open Data Catalog lists many datasets (e.g., disaster-awareness surveys) as CSV/XLSX — great that raw files exist (source: catalog.data.metro.tokyo.lg.jp).
  • Saitama and Hakodate clearly publish CSV portals (opendata.pref.saitama.lg.jp, harp.lg.jp) — adoption of CSV is common.
  • Niigata has a CSV manual (CSV creation guidance PDF) — signals awareness of machine-readability but docs aren't always enforced.

Look at this: notice.go.jp exposes notice CSVs directly (notice.go.jp/docs/status_notice.csv). That’s the kind of single-file access I like!

Common technical issues

  • Schema drift: column names, encodings (Shift_JIS vs UTF-8), and date formats differ between releases. That breaks ETL.
  • Metadata absence: many datasets lack DCAT/JSON-LD metadata to describe provenance, update frequency, and license.
  • API gap: some portals only expose bulk CSV downloads, not queryable APIs or pagination, which limits real-time apps.
  • PDF-first publications: some policy reports remain PDF-only, even when underlying tables could be CSV.

Policy vs data: KPI transparency

  • National programs like the Digital田園都市交付金 define KPIs (see guideline docs) and municipalities publish evaluation reports (e.g., Sukagawa). But time-series machine-readable KPI datasets are rare, making automated tracking and cross-municipality comparison hard.
  • 要するに、目標はあるけど、検証のためのデータがAPIで出てないケースが多い。

Concrete technical fixes (engineer-style!)

  • Enforce UTF-8 and RFC-compliant CSV (or provide JSON/JSONL). Use Niigata's CSV guidance as a baseline.
  • Publish a JSON Schema or CSV Schema (CSVW or data packages) per dataset so consumers can validate automatically.
  • Add a lightweight REST API (FastAPI) with OpenAPI so apps can query datasets without downloading whole files.
  • Provide DCAT metadata and an automated catalog index endpoint.
  • CI/CD for data: run schema validation, sample-row checks, and diffs on each dataset update.

Example: quick Python ingestion pattern

import pandas as pd

url = 'https://notice.go.jp/docs/status_notice.csv'

df = pd.read_csv(url, encoding='utf-8', parse_dates=['published_at'])

basic validation

assert 'id' in df.columns

print(df.head())

For schema validation, use pandera or jsonschema to assert types and ranges in CI.

Roadmap for municipalities (practical)

  • Start with CSVs + UTF-8 + header standardization.
  • Add a tiny API layer that serves JSON and CSV from the same canonical storage.
  • Publish machine-readable metadata (DCAT+JSON-LD) and example queries.
  • Link KPIs to time-series datasets and expose dashboards with exportable data.
  • まとめ

    Local governments already publish lots of useful data, but inconsistent formats, missing schemas, and limited APIs reduce impact. エンジニア的に言うと、フォーマットとメタデータを揃えてCIでチェックすれば、データはもっと再利用されます!

    おかむーから一言

    I've built teams and products on shaky data — trust me, small engineering hygiene wins huge civic outcomes. Let's push for CSV+schema+API across local govs and make policy measurable and reusable!