Code-Driven Manifesto: Evaluating Government Data and Systems from an Engineer's Lens

どうも〜おかむーです! Today I'm digging into how governments publish data, what engineers actually need, and where the biggest wins are for GovTech. Code-driven manifesto style — policy meets data and engineering.
- Public datasets exist, but formats and discoverability are all over the place
- Machine-readability (CSV/API/JSON) beats PDF every time — yet PDFs persist
- Small engineering fixes (APIs, schemas, CI, metadata) unlock huge civic value
結論
Government agencies often publish useful statistics and operational data (many go.jp CSV links exist!), but inconsistent formats, missing metadata, and PDF-first workflows block reuse. 要するに、データの“エンドポイント化”と品質ガバナンスをしてAPI + schema-firstで公開すれば、政策検証とサービス改善がぐっと進むということです。
Why I care (and you should, too)
エンジニア的に言うと、公開データがAPIで手元に落ちてくればダッシュボードも予算トラッキングも自動化できるんですよ。GovTech東京都の取り組み(govtechtokyo.or.jp)や日本の統計API(e-Stat)など、成功例はあるけど局所的なんです。
What I inspected (examples)
- Direct CSV endpoints: notice.go.jp/docs/status_nicter.csv, jinji.go.jp CSV, env.go.jp CSV — these exist and are great when present
- Agency portals and dashboards: GovTech Tokyo page on data utilization
- Domain naming quirks: international differences like .gov vs .gov.cn (.gov as TLD vs second-level) — small ops detail but matters for DNS/PKI and trust
これ見てくださいよ:CSVが直で置いてあるケースは解析がすごく早い。逆にPDFや画像に埋め込まれた表は、人力かOCR/Tabulaでの前処理が必要で非効率です。
Technical evaluation (what I look for)
I assessed typical government datasets against these axes:
- Availability: Is there a stable HTTP(S) endpoint? (good: /docs/status_nicter.csv)
- Machine-readability: CSV/JSON preferred; PDF is a blocker
- Metadata: schema description, field types, update timestamp, license
- API: REST/GraphQL or only page-based downloads?
- Provenance & versioning: changelogs, dataset versions
- Access controls and rate limits: any keys or throttles
Findings: many agencies publish CSVs (good!), but metadata is often missing or inconsistent. APIs exist (e.g., e-Stat) but many operational datasets are still static files or embedded in HTML/PDF.
Example: quick engineer workflow
If a ministry exposes CSV, you can start like this:
# fetch CSV
curl -sS https://notice.go.jp/docs/status_nicter.csv -o status_nicter.csv
quick inspect with Python
python - <<'PY'
import pandas as pd
df = pd.read_csv('status_nicter.csv')
print(df.dtypes)
print(df.head())
PY
要するに、API一本でこれをやれればデータパイプラインのオートメーションが進むんです。
PDF vs CSV — real pain points
- PDFs: layout changes break parsers. OCR/tabula pipelines are brittle. 要するに、PDFは人間向けであってマシン向けではない。
- CSV: no schema standard, encodings (Shift_JIS vs UTF-8) cause friction
- JSON/JSON-LD: ideal for nested structures and semantic clarity
Policy metrics: targets vs reality
Agencies publish targets in PDFs or reports, but operational metrics (monthly actuals) are often tucked into separate CSVs or not updated. That makes automated compliance checks or progress dashboards hard.
Example gap pattern:
- Policy doc: target = X by 2025 (PDF)
- Monthly performance: numbers in CSV but with different column names and no linkage to policy ID
Result: to compute progress you need brittle join logic and manual reconciliation.
Concrete technical recommendations
- Each dataset gets an ID, machine-readable metadata (schema, license, update cadence), and a canonical JSON/CSV/Parquet endpoint.
- Tech: use CKAN/Dataportal or a lightweight Git-backed registry with an index.json.
- Provide REST endpoints (or GraphQL) with OpenAPI specs. Developers love stable contracts.
- Example: /api/v1/policy_progress?policy_id=H123&from=2023-01
- Validate CSV/JSON schema on push (GitHub Actions) and emit dataset versions + changelog.
- Example: use goodtables or pandera in CI
- Policy metadata: id, target_value, target_date, unit
- Actuals: dataset includes policy_id to join reliably
Minimal OpenAPI idea (snippet)
openapi: 3.0.0
info:
title: Government Policy Progress API
paths:
/v1/policy/{id}/progress:
get:
parameters:
- name: id
in: path
required: true
responses:
'200':
content:
application/json:
schema:
$ref: '#/components/schemas/Progress'
components:
schemas:
Progress:
type: object
properties:
date:
type: string
value:
type: number
policy_id:
type: string
Implementation roadmap (practical steps)
- Week 0–4: inventory datasets, identify high-impact ones (budgets, procurement, health statistics)
- Week 4–12: add machine-readable metadata and stable endpoints, enable CORS
- Week 12–24: implement API wrappers, CI schema checks, publish OpenAPI
- Ongoing: community feedback loop, dataset SLAs, analytics on API usage
Use cases unlocked
- Real-time dashboards for policy performance
- Automated audits and KPI monitoring
- Civic apps and startups building on open data
- Research reproducibility and easier FOIA analyses
まとめ
Public data is a goldmine but often hidden behind PDFs and inconsistent files. エンジニア的に言うと、API-first + schema + CIで公開すれば、政策の可視化と検証が一気に現実的になります。小さな技術投資で透明性とアプリケーションエコシステムが激変するんですよ!
おかむーから一言
テクノロジーで社会をアップデートするのは本気で可能です。エンジニアリングの力で政策をもっと証拠ベースにしていきましょう!
Sources
- https://www.zhihu.com/question/40553450
- https://www.govtechtokyo.or.jp/services/data-utilization/
- https://www.zhihu.com/question/372341437
- https://note.govtechtokyo.jp/n/n77785a8254d6
- https://bus.gov.ru/
- https://www.zhihu.com/question/290714454
- https://metidx-gov.note.jp/n/n9468573c213b
- https://www.zhihu.com/question/6430289390
- https://www.trans-plus.jp/blog/column/202210_municipality-dx
- https://www.zhihu.com/question/38923279
- https://notice.go.jp/docs/status_nicter.csv
- https://www.jinji.go.jp/content/900024615.csv
- https://www.env.go.jp/content/900398071.csv
- https://www.inpit.go.jp/content/100869372.csv
- https://www.mhlw.go.jp/content/001429362.csv
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.