Code as Manifest: a Tech Audit of Government Data and Systems

IT Policy Proposals
Code as Manifest: a Tech Audit of Government Data and Systems

どうも〜おかむーです!Today I'm taking a stab at “コードで語るマニフェスト” — reading policy through the lens of data and engineering.

  • Government sites have pockets of machine-readable CSVs, but many KPIs live in PDFs — painful for reuse.
  • Where APIs exist (or raw CSVs), formats and metadata are inconsistent; that makes programmatic tracking of targets vs. outcomes hard.
  • Fixable by standardizing schemas, publishing APIs, and automating PDF→CSV pipelines — here's how, with code hints.

結論

Public data exists, but it's fragmented. 要するに、データは散在していて機械可読性がまちまち。エンジニア的に言うと、API-firstでスキーマを定義しないと政策の検証は再現できないんですよね。

Report: what I found and how I tested it

What the search results show (quick tour)

Look at this: Digital Agency (digital.go.jp) is the coordination hub, while specific program reports and KPI evaluations are often PDF-first (see the Digital田園都市構想 guideline / chisou.go.jp). Some ministries and local governments expose CSV files directly — notice.go.jp and several go.jp domains returned .csv URLs in the crawl — but there's no uniform API or metadata catalog like Data.gov or a consistent e-Stat interface for all datasets.

Why that matters

  • PDFs: human-readable but not machine-readable. 要するに、スクレイピングかOCRが必須ということです。
  • Raw CSV endpoints: great, but inconsistent column names, encodings, and no schema/versions.
  • APIs: rare and uneven. Where present, they aren't always documented or CORS-friendly.

Technical checks & code examples

Engineer-wise, here's a minimal pipeline to pull a CSV and compute KPI gap using Python:

import pandas as pd

url = 'https://example.go.jp/path/data.csv'

df = pd.read_csv(url, encoding='utf-8')

Normalize column names

df.columns = [c.strip().lower().replace(' ', '_') for c in df.columns]

Compute gap between target and actual

df['gap_pct'] = (df['actual'] - df['target']) / df['target'] * 100

print(df[['indicator', 'target', 'actual', 'gap_pct']].head())

And if you want to serve cleaned data as an API (FastAPI sketch):

from fastapi import FastAPI

import pandas as pd

app = FastAPI()

@api.get('/kpis')

def kpis():

df = pd.read_csv('cleaned_kpis.csv')

return df.to_dict(orient='records')

Dealing with PDFs

When the authoritative report is a PDF (chisou.go.jp examples), use tools like tabula-py or Camelot to extract tables, then validate via checksums and human review. Automate with CI that fails if table shapes change.

Schema & metadata

Propose a minimal JSON Schema for KPI tables: indicator_id, year, target_value, actual_value, unit, source_url, last_updated. Serve schema at /schema/kpi.json and publish a Data Package (frictionlessdata.io) manifest.

Policy gap analysis (how to compute)

  • Join program plans (targets) to periodic reports (actuals) by indicator_id and year.
  • Compute absolute and relative gaps, and visualize as small multiples.
  • Automate alerts when gap_pct > X% for >N consecutive periods.

改善提案(実務的)

  • API-first: every dataset should have a stable JSON/CSV endpoint + OpenAPI spec.
  • Central catalog: harvest metadata (DCAT or Data Package) across ministries into a searchable portal.
  • Versioned schemas: use JSON Schema and semantic versioning so downstream apps don't break.
  • PDF mitigation: require agencies to publish the underlying CSV/JSON when a PDF report is released.
  • Tooling: provide reference libraries (Python/R) and CI templates for extraction and validation.

まとめ

政策の数値検証は技術的に可能だけど、現状は手間が多い。これ見てくださいよ — 生データがある場所とない場所の差が大きすぎるんです。API, schema, automationの3点セットを導入すれば、政策の透明性と再現性が一気に上がります!

おかむーから一言

As an entrepreneur-engineer, I want governments to treat datasets like product APIs — discoverable, versioned, and reliable. Tech can make democracy more testable, so let's ship it!