Code-driven Manifesto: Making Municipal Data Actually Useful

どうも〜おかむーです! Today I want to do a slightly nerdy, slightly political post — in the spirit of「コードで語るマニフェスト」. エンジニア的に言うと、policy is only as verifiable as the data and APIs behind it. Let's dig in!
- This post inspects how Japanese governments publish data (PDF vs CSV vs API) and why machine-readability matters
- I reference gov resources (Digital Agency, e-Stat, Ministry rules) and show practical code patterns to make data useful
- I end with concrete engineering proposals so municipalities can turn policy claims into reproducible datasets
結論
Government policies often look great on paper, but from a technical perspective many promises are locked inside PDFs and isolated dashboards. 要するに、機械可読性(CSV/JSON/APIs)と versioned open endpoints が標準にならないと、政策の検証と民間活用は進まないということです。
Report
What I looked at (sources)
- Digital Agency / Digital Government materials including machine-readability guidance (see digital.go.jp documents on machine-readable rules)
- Ministry of Internal Affairs and Communications guidance on machine-readable statistics (soumu.go.jp)
- GovTech Tokyo projects and case studies showing dashboards and reuse patterns (govtechtokyo.or.jp)
- e-Stat (e-stat.go.jp) as the canonical national statistics API
これ見てくださいよ:最新のルール案では「ファイル形式は機械が直接読み取れる Excel や CSV 等となっているか」を求めている一方で、現場の公開は未だにPDFが多いんです!(参考: Digital Agency meeting notes)
Technical diagnosis
- PDF-first publishing: Many municipalities publish reports as PDF. PDFs are fine for humans but terrible for reproducible analysis. 要するに、スクレイピングかOCRしないと使えない。
- Incomplete APIs: Some central datasets (e-Stat) offer APIs, but local governments often lack standardized endpoints (no OpenAPI, no stable versioning).
- Format noise: Excel files with merged cells, embedded footnotes, and non-ASCII encodings break automated pipelines.
- Metrics vs time series: Policy targets are often stated without machine-readable time series, so measuring progress programmatically is hard.
Example: How an engineer would approach it
1) If data is CSV/Excel with clean headers: pandas does the job in two lines
import pandas as pd
df = pd.read_csv('https://example.gov/data/municipal_kpi.csv')
df['date'] = pd.to_datetime(df['date'])
2) If only PDF is available: use tabula or Camelot to extract tables, but expect cleanup
tabula -p all -f CSV 2025_policy_report.pdf -o extracted.csv
3) For APIs: prefer JSON with schema and pagination. Example e-Stat style request (pseudo):
import requests
resp = requests.get('https://api.e-stat.go.jp/rest/2.1/app/getStatsData', params={
'appId': 'YOUR_KEY', 'statsDataId':'000345'} )
data = resp.json()
Policy targets vs reality — an engineering check
- Policy: "Reduce waiting time X by 30% by 2026" — often stated in press releases (PDF).
- What you need to verify: monthly time-series of waiting times, metadata describing measurement method, and a stable API endpoint.
- Common gap: municipalities publish an annual PDF with an aggregate number. That prevents trend analysis and statistical testing.
Practical improvements (concrete)
- Publish primary indicators as CSV/JSON time series, not only PDF summaries. Follow the Digital Agency's machine-readable guidance.
- Provide an OpenAPI spec and simple API tokens (or open endpoints for public data). That enables developers to build reproducible checks and dashboards.
- Version datasets (date-stamped S3 buckets or GitHub releases) so historical corrections are trackable.
- Normalize encodings (UTF-8), avoid merged cells in spreadsheets, and publish a JSON Schema for each dataset.
- Create a minimal example repo: CKAN or simple static hosting with index.json describing datasets. GovTech Tokyo projects show how dashboards can be reused when data is clean.
Low-friction migration plan for a municipality
Example index.json snippet (concept):
{
"datasets": [
{"id":"waiting_time","title":"Clinic waiting times","format":"csv","url":"https://data.city.example.jp/waiting_time.csv","schema_url":"https://.../waiting_time.schema.json"}
]
}
まとめ
- 現状: PDFや非標準Excelが多く、検証が難しい。APIは部分的で統一されていない。
- 要るもの: CSV/JSON time-series、OpenAPI、versioning、JSON Schema。
- 効果: 政策の透明性が上がり、市民・事業者がエビデンスに基づいてサービスや検証を行えるようになる。
おかむーから一言
テクノロジーで社会をアップデートするって言ってるんだから、まずはデータをちゃんと開けようよ!エンジニアとしても市民としても、検証可能な自治体こそ信頼できると思うんです。迅速な改善、期待してます!
Sources
- https://ja.wikipedia.org/wiki/%E5%85%AC%E5%85%B1
- https://www.intec.co.jp/column/smartcity-08.html
- https://kotobank.jp/word/%E5%85%AC%E5%85%B1-494676
- https://www.digital.go.jp/resources/data_case_study_private
- https://lifeap.co.jp/column/2026/03/18/understanding-public-facilities-definition-types-examples-usage-and-management/
- https://www.zhihu.com/question/290714454
- https://www.digital.go.jp/assets/contents/node/basic_page/field_ref_resources/256dcba6-b936-4031-b88d-3abb27e27f9b/f7af0ca4/20260331_meeting_executive_outline_06.pdf
- https://www.zhihu.com/question/6430289390
- https://www.soumu.go.jp/menu_news/s-news/01toukatsu01_02000186.html
- https://www.zhihu.com/question/38923279
- https://support.yahoo-net.jp/voc/s/ytop-sp
- https://www.govtechtokyo.or.jp/services/data-utilization/
- https://support.yahoo-net.jp/PccTop/s/topic/0TO2r000000GnthGAC/yahoo-japan%E3%83%88%E3%83%83%E3%83%97%E3%83%9A%E3%83%BC%E3%82%B8%E5%85%A8%E8%88%AC
- https://note.govtechtokyo.jp/n/n77785a8254d6
- https://support.yahoo-net.jp/PccHelpcenter/s/
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.