Code Manifest: Reading Japan's Open Data Like an Engineer

IT Policy Proposals
Code Manifest: Reading Japan's Open Data Like an Engineer

どうも〜おかむーです!今日はちょっとエンジニアっぽい話をしますよ〜

  • Public datasets exist across Tokyo, Niigata, Saitama, Hakodate and national portals, but formats and freshness vary a lot
  • Machine-readability (CSV/JSON/API) often mixed with PDFs and legacy encodings — so automated pipelines break
  • Practical fixes: publish stable APIs, DCAT metadata, UTF-8 CSV/JSON, schema validation and CI for data

結論

エンジニア的に言うと、政策をコードで語るには「データがAPIで、スキーマが明確で、更新が信頼できる」ことが必須なんですよね。現状は部分的に達成されているけど、安定した運用と相互運用性が足りない。要するに、データ公開の仕組みをエンジニアリングで固めれば、政策評価がずっと早く、正確になるということです。

Technical inventory: what I found

Public catalogs and formats

これ見てくださいよ:Tokyo’s open data catalog (https://catalog.data.metro.tokyo.lg.jp/dataset) exposes datasets (CSV/XLSX) including disaster-awareness surveys. Niigata publishes CSVs and even a CSV manual (https://www.city.niigata.lg.jp/.../csv_manual_v1.1.pdf) — nice! Saitama and Hakodate list many CSVs via their portals. Nationally, Data StaRt (stat.go.jp) provides use-cases and soumu/go.jp hosts CSV exports.

  • Strengths: many datasets are available as CSV/Excel and sometimes have explicit open licenses (Hakodate noted CC-BY)
  • Weaknesses: mixed encodings, occasional PDF-only releases, stale snapshots (e.g., Niigata population data referencing 平成22年 / 2010), inconsistent metadata and no uniform API spec

Machine-readability and APIs

APIs are patchy. Some portals only allow manual CSV download; others support dataset lists but lack RESTful JSON endpoints with pagination, CORS and schema discovery. エンジニア的に言うと、API一本で解決する話なんですよね — stable endpoints + machine-readable metadata = huge developer productivity gains.

Example: quick ingestion pattern

Here’s a minimal Python snippet to read a CSV from a public URL and validate columns (pseudocode-like, quick start):

import pandas as pd

url = 'https://catalog.data.metro.tokyo.lg.jp/dataset/xxx.csv'

df = pd.read_csv(url, encoding='utf-8')

basic validation

required = ['prefecture', 'ward', 'population', 'year']

missing = [c for c in required if c not in df.columns]

if missing:

raise ValueError('Missing columns: ' + ','.join(missing))

convert types

df['year'] = df['year'].astype(int)

要するに、安定して自動で読めることが大事なんです。

Policy vs data: where numbers fall short

  • Temporal gap: datasets used for planning (e.g., regional population or disaster awareness) are sometimes based on decade-old snapshots — that weakens policy feedback loops
  • KPI transparency: policy targets often lack machine-readable time-series that show progress. If you can't programmatically pull quarterly indicators, you can't build automated dashboards or alerts

Concrete engineering improvements

  • Publish DCAT-compliant metadata and machine-readable licenses (JSON-LD)
  • Offer REST APIs with pagination, CORS, and stable versioning; provide bulk dumps for reproducibility
  • Standardize CSV conventions: UTF-8, RFC4180, ISO date formats, clear column names and units (or better: JSON Schema)
  • Implement data CI: schema checks, value ranges, freshness tests in CI pipelines (GitHub Actions etc.)
  • Provide geospatial outputs as GeoJSON / WFS for map apps
  • Ship example clients and Python/R notebooks demonstrating usage
  • Implementation sketch (stack)

    • Portal: CKAN or DataHub with DCAT export
    • Schema registry: JSON Schema + pandera for pandas validation
    • ETL: Airflow or GitHub Actions for scheduled imports and tests
    • API layer: FastAPI serving JSON endpoints with OpenAPI docs

    まとめ

    現場のデータはあるけど、エンジニア目線での「安定して読み取り可能」な形に整えることが最短の改善ルートです。CSVだけで完結しないで、API・メタデータ・CIをセットで導入するのが王道。これで政策検証がグッと速くなるはず!

    おかむーから一言

    Techで社会をアップデートするのが俺のミッション。データの細かいところを直せば、政策の精度と市民の信頼が一気に上がりますよ〜一緒にやりましょう!