Code-Driven Manifest: Auditing Government Data and Systems from an Engineer's Lens

IT Policy Proposals
Code-Driven Manifest: Auditing Government Data and Systems from an Engineer's Lens

どうも〜おかむーです! Hi — today I want to take a technical look at how Japanese national and municipal data/systems are published and how we can make them actually useful for policy and developers.

  • Many government reports are published as PDFs; machine-readability is low and APIs are sparse
  • Some good initiatives exist (Digital Agency, GovTech Tokyo) but operational gaps remain (KPI wiring, data formats)
  • Concrete engineering fixes: CSV/JSON-first publishing, OpenAPI, canonical identifiers, and reproducible ETL

結論

Public data availability is improving thanks to digital.go.jp and GovTech efforts, but the day-to-day reality is still: a lot of PDF, inconsistent schemas, and missing APIs. Engineers can turn policy intent into measurable impact by demanding machine-readable outputs, standardized identifiers, and stable APIs — that's where the manifesto becomes code.

Report

What I looked at (sources)

  • Digital Agency overview: https://www.digital.go.jp/ — sets the national DX direction
  • GovTech Tokyo data services & dashboards: https://www.govtechtokyo.or.jp/services/data-utilization/ and related writeups (note.govtechtokyo.jp)
  • Policy KPI guidance and evaluations (example: Digital田園都市交付金 evaluation guideline PDF on chisou.go.jp and Sukagawa city evaluation page) https://www.chisou.go.jp/sousei/pdf/r5_guideline-checkaction.pdf https://www.city.sukagawa.fukushima.jp
  • e-Stat and data.go.jp as canonical open-data portals

これ見てくださいよ — a lot of municipal evaluation docs and funding reports are PDF-only. That means numbers exist but are locked away from programmatic use. Engineers hate that, because reproducibility and automation become impossible.

Key technical problems

  • PDF-first publishing
  • - Problem: tables inside PDFs are not machine-readable and often require manual extraction (tabula, OCR).

    - Consequence: delayed insights, human errors, poor reproducibility.

  • Inconsistent schemas across municipalities
  • - No canonical column names or identifiers (use of different date formats, different entity IDs) hampers merging datasets.

  • Sparse or undocumented APIs
  • - Some portals (e-Stat) have APIs, but many local governments only provide HTML/PDF downloads without OpenAPI specs.

  • KPI vs. measured reality disconnect
  • - Policy docs list KPIs (e.g., Digital田園都市 targets) but progress reports are buried in PDFs or non-standard tables, making trend analysis hard.

    Concrete engineering fixes (practical)

    • Publish CSV/JSON as primary artifacts alongside human-friendly PDFs. Use data.go.jp and link exact file URLs in policy pages.
    • Provide an OpenAPI (Swagger) spec for any REST endpoints. Example minimal OpenAPI snippet (pseudo):
    openapi: 3.0.0
    

    info:

    title: Municipal KPI API

    version: 1.0.0

    paths:

    /kpis:

    get:

    parameters:

    - name: year

    in: query

    schema:

    type: integer

    responses:

    '200':

    content:

    application/json:

    schema:

    type: array

    • Use canonical identifiers: corporate numbers, JIS codes for municipalities, ISO date formats (YYYY-MM-DD).
    • Publish change-logs and dataset versioning (semantic versioning of datasets) so time series are reproducible.
    • Provide example code snippets and reproducible notebooks. Example ETL pattern in Python:
    import pandas as pd
    

    ingest CSV from data portal

    df = pd.read_csv('https://data.example.jp/municipal_kpis.csv')

    normalize columns

    df.columns = df.columns.str.strip().str.lower()

    df['date'] = pd.to_datetime(df['date'])

    group and compare target vs actual

    summary = df.groupby('municipality').agg({'target':'sum','actual':'sum'})

    • Where only PDFs exist, include machine-readable attachments or extractable CSVs using tools like tabula-py and publish the transformed output with provenance.

    Policy measurement: wiring KPIs to dashboards

    • Use GovTech Tokyo's dashboard approach as reference: combine centralized dashboards for national-level KPIs and federated data feeds from municipalities.
    • Automate KPI ingestion pipelines (cron + CI) with unit tests that assert schema and ranges — fail loudly when data breaks.

    まとめ

    Public-sector digitalization is more than a slogan: it's about turning PDFs into streams, policies into metrics, and ambiguity into schemas. The tech stack is simple: CSV/JSON-first, OpenAPI, canonical IDs, reproducible ETL and public dashboards. Do that and policy evaluation stops being manual grunt work and becomes continuous, auditable feedback.

    おかむーから一言

    I built startups and shipped full-stack GovTech projects — so I get the messy reality. Let's stop treating policy reports as ceremonial PDFs and start shipping data as code. That's how we actually change things!