Code the Manifest: Reading Local Gov Data as an Engineer

IT Policy Proposals
Code the Manifest: Reading Local Gov Data as an Engineer

どうも〜おかむーです!Today I'll take a technical look at how Japanese local governments publish data — and what engineers can do with it.

  • Tokyo, Niigata, Saitama and many cities publish useful CSVs, but metadata and APIs are inconsistent
  • Machine-readability (CSV vs PDF), missing schema and undocumented fields block reuse
  • A few practical fixes (DCAT, CKAN APIs, schema, ETL examples) unlock much higher public value

結論

Municipal open data is there, but it's fragmented and often under-specified. エンジニア的に言うと、データをAPIとスキーマで正しく公開すれば、政策の検証とサービス化は一気に進むということです。

Report: what I found and how to fix it

これ見てくださいよ — Tokyo Open Data Catalog (https://catalog.data.metro.tokyo.lg.jp/dataset) hosts a historical dataset:都営水道の水源量 since 1899 in CSV form. Amazing! But many other portals (Niigata, Saitama, Yokohama) either have CSVs with no descriptions or rely on CKAN instances with minimal metadata. That means data scientists waste time guessing column meanings rather than analyzing policy impact.

What I checked

  • Data formats: CSVs exist, but some datasets are only in PDF or lack field descriptions (see Yokohama entry noting “このデータセットには説明がありません”). PDFs are not machine-readable —要するに使いづらいということです。
  • APIs: Some municipalities run CKAN (e.g., Daisen's CKAN) which provides APIs, but API usage and authentication vary. No uniform approach across municipalities.
  • Metadata: DCAT/JSON-LD metadata rarely present. No standardized schema for similar datasets (population, healthcare, water supply).

Technical examples

  • Quick retrieval (Python + pandas):
import pandas as pd

url = 'https://catalog.data.metro.tokyo.lg.jp/dataset/xxx.csv' # replace with actual dataset URL

df = pd.read_csv(url)

print(df.head())

  • Validation hint: use pandera or jsonschema to define expected columns and types so downstream dashboards don't silently break.

Policy gap analysis

  • Tokyo water dataset gives long-run supply volumes. If you overlay that with population and consumption (Niigata / Saitama population CSVs), you can detect structural stress periods. Right now, lack of consistent timestamps and unit metadata prevents automated comparison — so policy targets vs real trends can't be easily validated.
  • Example: policy target = reduce per capita consumption by 10% in 5 years. To verify, you need per-capita water use time series with stable unit annotations. Many datasets lack that.

Improvement roadmap (practical)

  • Publish DCAT metadata and machine-readable license for every dataset (helps catalogs and search)
  • Standardize schemas for common domains (population, care facilities, water supply) and publish JSON Schema examples
  • Expose CKAN/REST APIs with sample queries and CORS enabled so frontend apps can fetch directly
  • Provide small example notebooks (Colab/Gist) demonstrating ETL and common joins (water x population)
  • Add automated validation on publish (CI that runs schema checks and basic stats)
  • Open-data use cases unlocked

    • Real-time dashboards combining river/IoT sensors and historical supply data to detect drought risk
    • Civic apps that geolocate care facilities (Saitama dataset) and offer booking/alerts
    • Machine-learned forecasts for water demand to guide procurement and conservation policy

    まとめ

    地方自治体はデータ公開をしているけど、エンジニア視点だと“使えるデータ”にするためのメタデータ、API、スキーマが足りない。要するに、いまの公開スタイルだと再利用にコストがかかりすぎる。だからこそ、DCAT、CKAN API、JSON Schemaといった小さな改善で政策の検証力は一気に上がりますよね!

    おかむーから一言

    データをちゃんと公開すれば、政策はコードで検証できるんです。技術で社会をアップデートするって、まさにこういう地道な改善から始まります!