Code as Manifest: a Tech Audit of Government Data and Systems

どうも〜おかむーです!Today I'm taking a stab at “コードで語るマニフェスト” — reading policy through the lens of data and engineering.
- Government sites have pockets of machine-readable CSVs, but many KPIs live in PDFs — painful for reuse.
- Where APIs exist (or raw CSVs), formats and metadata are inconsistent; that makes programmatic tracking of targets vs. outcomes hard.
- Fixable by standardizing schemas, publishing APIs, and automating PDF→CSV pipelines — here's how, with code hints.
結論
Public data exists, but it's fragmented. 要するに、データは散在していて機械可読性がまちまち。エンジニア的に言うと、API-firstでスキーマを定義しないと政策の検証は再現できないんですよね。
Report: what I found and how I tested it
What the search results show (quick tour)
Look at this: Digital Agency (digital.go.jp) is the coordination hub, while specific program reports and KPI evaluations are often PDF-first (see the Digital田園都市構想 guideline / chisou.go.jp). Some ministries and local governments expose CSV files directly — notice.go.jp and several go.jp domains returned .csv URLs in the crawl — but there's no uniform API or metadata catalog like Data.gov or a consistent e-Stat interface for all datasets.
Why that matters
- PDFs: human-readable but not machine-readable. 要するに、スクレイピングかOCRが必須ということです。
- Raw CSV endpoints: great, but inconsistent column names, encodings, and no schema/versions.
- APIs: rare and uneven. Where present, they aren't always documented or CORS-friendly.
Technical checks & code examples
Engineer-wise, here's a minimal pipeline to pull a CSV and compute KPI gap using Python:
import pandas as pd
url = 'https://example.go.jp/path/data.csv'
df = pd.read_csv(url, encoding='utf-8')
Normalize column names
df.columns = [c.strip().lower().replace(' ', '_') for c in df.columns]
Compute gap between target and actual
df['gap_pct'] = (df['actual'] - df['target']) / df['target'] * 100
print(df[['indicator', 'target', 'actual', 'gap_pct']].head())
And if you want to serve cleaned data as an API (FastAPI sketch):
from fastapi import FastAPI
import pandas as pd
app = FastAPI()
@api.get('/kpis')
def kpis():
df = pd.read_csv('cleaned_kpis.csv')
return df.to_dict(orient='records')
Dealing with PDFs
When the authoritative report is a PDF (chisou.go.jp examples), use tools like tabula-py or Camelot to extract tables, then validate via checksums and human review. Automate with CI that fails if table shapes change.
Schema & metadata
Propose a minimal JSON Schema for KPI tables: indicator_id, year, target_value, actual_value, unit, source_url, last_updated. Serve schema at /schema/kpi.json and publish a Data Package (frictionlessdata.io) manifest.
Policy gap analysis (how to compute)
- Join program plans (targets) to periodic reports (actuals) by indicator_id and year.
- Compute absolute and relative gaps, and visualize as small multiples.
- Automate alerts when gap_pct > X% for >N consecutive periods.
改善提案(実務的)
- API-first: every dataset should have a stable JSON/CSV endpoint + OpenAPI spec.
- Central catalog: harvest metadata (DCAT or Data Package) across ministries into a searchable portal.
- Versioned schemas: use JSON Schema and semantic versioning so downstream apps don't break.
- PDF mitigation: require agencies to publish the underlying CSV/JSON when a PDF report is released.
- Tooling: provide reference libraries (Python/R) and CI templates for extraction and validation.
まとめ
政策の数値検証は技術的に可能だけど、現状は手間が多い。これ見てくださいよ — 生データがある場所とない場所の差が大きすぎるんです。API, schema, automationの3点セットを導入すれば、政策の透明性と再現性が一気に上がります!
おかむーから一言
As an entrepreneur-engineer, I want governments to treat datasets like product APIs — discoverable, versioned, and reliable. Tech can make democracy more testable, so let's ship it!
Sources
- https://www.zhihu.com/question/290714454
- https://metidx-gov.note.jp/n/n9468573c213b
- https://www.zhihu.com/question/6430289390
- https://zenn.dev/govtechtokyo/articles/b65dc687e50918
- https://www.zhihu.com/question/38923279
- https://ja.wikipedia.org/wiki/%E3%83%87%E3%82%B8%E3%82%BF%E3%83%AB
- https://www.chisou.go.jp/sousei/pdf/r5_guideline-checkaction.pdf
- https://column.nippoukun.bpsinc.jp/what-is-digitization/
- https://www.city.sukagawa.fukushima.jp/shisei/gyoseiunei/keikaku/chiho_sosei/1015604/4045.html
- https://www.digital.go.jp/
- https://notice.go.jp/docs/status_nicter.csv
- https://www.jinji.go.jp/content/900024615.csv
- https://www.env.go.jp/content/900398071.csv
- https://www.inpit.go.jp/content/100869372.csv
- https://www.mhlw.go.jp/content/001429362.csv
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.