Code-driven Manifesto: Making Municipal Data Actually Useful

IT Policy Proposals
Code-driven Manifesto: Making Municipal Data Actually Useful

どうも〜おかむーです! Today I want to do a slightly nerdy, slightly political post — in the spirit of「コードで語るマニフェスト」. エンジニア的に言うと、policy is only as verifiable as the data and APIs behind it. Let's dig in!

  • This post inspects how Japanese governments publish data (PDF vs CSV vs API) and why machine-readability matters
  • I reference gov resources (Digital Agency, e-Stat, Ministry rules) and show practical code patterns to make data useful
  • I end with concrete engineering proposals so municipalities can turn policy claims into reproducible datasets

結論

Government policies often look great on paper, but from a technical perspective many promises are locked inside PDFs and isolated dashboards. 要するに、機械可読性(CSV/JSON/APIs)と versioned open endpoints が標準にならないと、政策の検証と民間活用は進まないということです。

Report

What I looked at (sources)

  • Digital Agency / Digital Government materials including machine-readability guidance (see digital.go.jp documents on machine-readable rules)
  • Ministry of Internal Affairs and Communications guidance on machine-readable statistics (soumu.go.jp)
  • GovTech Tokyo projects and case studies showing dashboards and reuse patterns (govtechtokyo.or.jp)
  • e-Stat (e-stat.go.jp) as the canonical national statistics API

これ見てくださいよ:最新のルール案では「ファイル形式は機械が直接読み取れる Excel や CSV 等となっているか」を求めている一方で、現場の公開は未だにPDFが多いんです!(参考: Digital Agency meeting notes)

Technical diagnosis

  • PDF-first publishing: Many municipalities publish reports as PDF. PDFs are fine for humans but terrible for reproducible analysis. 要するに、スクレイピングかOCRしないと使えない。
  • Incomplete APIs: Some central datasets (e-Stat) offer APIs, but local governments often lack standardized endpoints (no OpenAPI, no stable versioning).
  • Format noise: Excel files with merged cells, embedded footnotes, and non-ASCII encodings break automated pipelines.
  • Metrics vs time series: Policy targets are often stated without machine-readable time series, so measuring progress programmatically is hard.

Example: How an engineer would approach it

1) If data is CSV/Excel with clean headers: pandas does the job in two lines

import pandas as pd

df = pd.read_csv('https://example.gov/data/municipal_kpi.csv')

df['date'] = pd.to_datetime(df['date'])

2) If only PDF is available: use tabula or Camelot to extract tables, but expect cleanup

tabula -p all -f CSV 2025_policy_report.pdf -o extracted.csv

3) For APIs: prefer JSON with schema and pagination. Example e-Stat style request (pseudo):

import requests

resp = requests.get('https://api.e-stat.go.jp/rest/2.1/app/getStatsData', params={

'appId': 'YOUR_KEY', 'statsDataId':'000345'} )

data = resp.json()

Policy targets vs reality — an engineering check

  • Policy: "Reduce waiting time X by 30% by 2026" — often stated in press releases (PDF).
  • What you need to verify: monthly time-series of waiting times, metadata describing measurement method, and a stable API endpoint.
  • Common gap: municipalities publish an annual PDF with an aggregate number. That prevents trend analysis and statistical testing.

Practical improvements (concrete)

  • Publish primary indicators as CSV/JSON time series, not only PDF summaries. Follow the Digital Agency's machine-readable guidance.
  • Provide an OpenAPI spec and simple API tokens (or open endpoints for public data). That enables developers to build reproducible checks and dashboards.
  • Version datasets (date-stamped S3 buckets or GitHub releases) so historical corrections are trackable.
  • Normalize encodings (UTF-8), avoid merged cells in spreadsheets, and publish a JSON Schema for each dataset.
  • Create a minimal example repo: CKAN or simple static hosting with index.json describing datasets. GovTech Tokyo projects show how dashboards can be reused when data is clean.

Low-friction migration plan for a municipality

  • Audit existing publications: list PDF tables that should be time-series.
  • Convert authoritative tables to CSV and publish to a stable URL with an index.json (example schema provided below).
  • Add a simple read-only API backed by the same CSVs and an OpenAPI spec.
  • Work with civic tech groups (e.g., GovTech Tokyo) for user testing and discoverability.
  • Example index.json snippet (concept):

    {
    

    "datasets": [

    {"id":"waiting_time","title":"Clinic waiting times","format":"csv","url":"https://data.city.example.jp/waiting_time.csv","schema_url":"https://.../waiting_time.schema.json"}

    ]

    }

    まとめ

    • 現状: PDFや非標準Excelが多く、検証が難しい。APIは部分的で統一されていない。
    • 要るもの: CSV/JSON time-series、OpenAPI、versioning、JSON Schema。
    • 効果: 政策の透明性が上がり、市民・事業者がエビデンスに基づいてサービスや検証を行えるようになる。

    おかむーから一言

    テクノロジーで社会をアップデートするって言ってるんだから、まずはデータをちゃんと開けようよ!エンジニアとしても市民としても、検証可能な自治体こそ信頼できると思うんです。迅速な改善、期待してます!