Code-Savvy Manifesto: Auditing Government Data and Systems

どうも〜おかむーです!今日は政府・自治体のデータ公開をコードの目で検証していきますよ〜
- Government CSV endpoints exist (e.g. notice.go.jp/status_nicter.csv) but machine-readability and API support are inconsistent
- PDFs and ad-hoc CSVs hurt reuse; a small engineering push (OpenAPI + JSON-LD + CSVW) would unlock lots of value
- Practical fixes: stable URLs, schema, CORS, rate limits, and a simple ETL + dashboard pipeline
結論
Public-sector data is often available, but usability is the bottleneck. エンジニア的に言うと、データが“存在している”だけで“使える”状態になっていないことが多いんです。CSVを置くのは良いけど、スキーマが不明確、バージョン管理がない、CORSやAPIがない、これが課題です。要するに、データ公開は技術的な製品設計が必要ということです。
レポート本文
現状の観察(データソース例)
これ見てくださいよ:notice.go.jpのCSV(/docs/status_nicter.csv)や、各省の公開CSV(jinji.go.jp/content/900024615.csv、env.go.jp/content/900398071.csv、mhlw.go.jp/content/001429362.csv)がある。GovTech東京の取り組み(https://www.govtechtokyo.or.jp/services/data-utilization/)もダッシュボード作りを進めている。
しかし問題点は明確:
- フォーマットのばらつき(CSVのエンコーディング、カラム名、不統一な日時)
- APIの欠如(毎回ファイルを叩くには限界がある)
- メタデータ不足(スキーマ、更新履歴、ライセンス情報が散逸)
技術的検証/サンプルコード
エンジニア的に言うと、API一本で済む話なんですよね。まずは生データを取り込む簡単な例:
import pandas as pd
url = 'https://notice.go.jp/docs/status_nicter.csv'
df = pd.read_csv(url)
print(df.head())
スキーマ検証はjsonschemaやpanderaで自動化できる。CSVWやJSON-LDでメタデータを添えれば、データカタログとの連携も楽になります。
政策評価のやり方(数値目標と実績の比較)
政策は目標と実績を時系列で照らし合わせるのが鉄板。やり方:
- 目標値を機械可読に(JSON/CSV)
- 実績データをETLで統一スキーマに変換
- ダッシュボード(Grafana/Metabase)で差分とトレンドを公開
これで市民も、議会も、エンジニアも検証できるようになります。
改善提案(短期〜中期)
- 全公開データにCSVW/JSON-LDメタデータを付与
- OpenAPI + JSON endpoints for common datasets
- Stable URLs + semantic versioning for datasets
- CORS enabled + basic rate limits and API keys
- Provide example notebooks (Jupyter) and schema tests
まとめ
技術的には、今すぐできる改善が山ほどあります。PDFのまま放置せず、CSVだけでも標準化して、APIを一本公開するだけで利活用は飛躍的に上がるんですよ!
おかむーから一言
I've built startups and shipped GOV tech stacks — the gap isn't money, it's product thinking. Make datasets first-class products: version, document, and ship APIs. Let's iterate fast and open government data for real.
Sources
- https://www.zhihu.com/question/40553450
- https://www.govtechtokyo.or.jp/services/data-utilization/
- https://www.zhihu.com/question/372341437
- https://note.govtechtokyo.jp/n/n77785a8254d6
- https://www.zhihu.com/tardis/zm/art/1924492115896960699
- https://www.zhihu.com/question/290714454
- https://metidx-gov.note.jp/n/n9468573c213b
- https://www.zhihu.com/question/6430289390
- https://www.trans-plus.jp/blog/column/202210_municipality-dx
- https://www.zhihu.com/question/38923279
- https://notice.go.jp/docs/status_nicter.csv
- https://www.jinji.go.jp/content/900024615.csv
- https://www.env.go.jp/content/900398071.csv
- https://www.inpit.go.jp/content/100869372.csv
- https://www.mhlw.go.jp/content/001429362.csv
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.