Code-Driven Manifesto: Diagnosing Japan’s Gov Data Plumbing

IT Policy Proposals
Code-Driven Manifesto: Diagnosing Japan’s Gov Data Plumbing

Hey — Okamu here! Today we're doing a little GovTech autopsy with a coder's scalpel.

  • Governments publish policies and promises, but often bury measurements in PDFs or scattered APIs.
  • Machine-readable rules (CSV/Excel/JSON) and a single API contract would unlock reuse and accountability.
  • Small engineering fixes (API-first, schema, ETL pipelines) give outsized impact for transparency and policy evaluation.

結論

The policy is often fine on paper, but the data plumbing is the bottleneck. Japan has an API catalog (e-Gov) and an open-data push from Soumu/Digital Agency, yet many statistics remain PDF-locked or inconsistently formatted. Engineer-wise, this is solvable: standardize machine-readable outputs (CSV/JSON), publish OpenAPI specs, and provide canonical ETL snippets so both auditors and startups can reproduce performance metrics.

Report

What the sources say (quick)

  • e-Gov portal lists administrative APIs and an API catalog (https://api-catalog.e-gov.go.jp/ ; https://www.e-gov.go.jp/digital-government/api) — good sign!
  • Soumu's open data strategy emphasizes standard data models and shared APIs (soumu.go.jp)
  • Digital Agency draft rules on machine-readability explicitly rank formats and push CSV/Excel over PDFs (digital.go.jp PDF)

These are the right policies, but implementation gaps remain: datasets still arrive as PDFs, tables aren’t standardized, and API coverage is uneven.

Technical diagnosis

  • Machine readability
  • - Problem: PDFs and scanned documents are still common for release notes and statistical tables. The draft rule classifies formats by "Level 1.." and recommends CSV/Excel for machine-readability. In practice, many "open" datasets are effectively locked.

    - Impact: Inhibits reproducible analysis and real-time dashboards.

  • API hodgepodge
  • - Problem: e-Gov API catalog lists many endpoints, but no single canonical data model or versioning guideline. Different ministries expose similar concepts (population, fiscal) with different field names and encodings.

    - Impact: Integration cost for civic apps skyrocket.

  • Provenance & performance metrics
  • - Problem: Policy targets (e.g., open-data conversion goals) are set, but outcome data often lacks time-series machine-readable exports to validate progress.

    - Impact: Hard to audit policy effectiveness programmatically.

    Concrete engineering fixes

    • API-first with OpenAPI: require new datasets to publish an OpenAPI/Swagger file plus example JSON/CSV. That allows client codegen and schema validation.
    • CSV-first publication: always publish CSV/Excel alongside PDFs. Where PDFs are unavoidable, publish extracted CSVs and the extraction script.
    • Canonical vocabularies: publish a shared JSON-LD context or simple data dictionary for common domains (population, budgets, licenses).
    • Versioning & metadata: use datapackage.json or DCAT metadata with checksum, publish last-updated and schema version.

    Example: tiny reproducible ETL (Python)

    # fetch a CSV from a hypothetical e-Gov endpoint and compute monthly deltas
    

    import requests, pandas as pd

    url = 'https://api.example.gov.jp/v1/population.csv'

    df = pd.read_csv(url)

    df['date'] = pd.to_datetime(df['year_month'])

    df = df.sort_values('date')

    df['delta'] = df['value'].diff()

    print(df.tail())

    Engineer-wise, this is the minimum reproducible building block. If the endpoint instead returns PDF, you'd add a step with tabula-py or OCR, plus checksum and manual review.

    Policy vs Implementation: a quick gap analysis

    • Policy: Soumu’s three-pillars open data push (experiments, industry collaboration, agency-held data) is sound.
    • Implementation gap: Without mandated machine-readable outputs and API contracts, these pillars are weak. The draft machine-readability rules are promising, but need enforcement and developer-friendly examples.

    まとめ

    Check this out: the pieces are there — e-Gov APIs, Soumu strategy, Digital Agency rules. What’s missing is consistent engineering discipline: API-first, CSV/JSON by default, shared schemas, and reproducible ETL examples. Do that and startups, researchers, and civic hackers can turn policy into verifiable impact.

    おかむーから一言

    I’ve built products from zero to scale — the gov stack just needs the same engineering hygiene. Small rules (OpenAPI, CSV-first) unlock huge civic value. Let’s ship the plumbing, then judge the policy!