Manifest in Code: Evaluating Japan's Administrative Data Practices from an Engineering Lens

IT Policy Proposals
Manifest in Code: Evaluating Japan's Administrative Data Practices from an Engineering Lens

どうも〜おかむーです! Today I'm digging into government and municipal data practices with a coder's eye — "code speaks the manifesto" style. I'll look at machine‑readability rules, system standardization, and realistic engineering fixes you can actually ship.

  • Japan's Digital Agency has formalized machine‑readability rules (Mar 31, 2026) but many datasets remain trapped in PDF.
  • Local government system standardization and gov‑cloud migration are underway, but data APIs, schemas, and CI for data quality are still immature.
  • Concrete engineering fixes (schema registry, API-first publishing, CI tests, reference libs) would turn policy into usable data.

結論

The policy direction is right — requiring Excel/CSV and pushing system standardization (see Digital Agency and Cabinet Secretariat materials) — but execution is where it breaks: too much PDF, too few stable APIs, and no shared schema governance. 要するに、政策は正しいけど“データデリバリーパイプライン”が整っていないということです。

Report

Where the data is now

Look at this: Digital Agency's docs on machine‑readability (https://www.digital.go.jp) decide formats and levels; Cabinet Secretariat summarizes decisions (https://www.cas.go.jp). The Ministry of Internal Affairs provides guidance on local system standardization (https://www.soumu.go.jp). These are solid governance signals, but the actual published artifacts from many municipalities are still PDFs or siloed Excel files.

Why that matters (engineer view):

  • PDF is non‑deterministic for scraping and breaks data contracts.
  • Excel sheets differ per municipality => schema drift and brittle ETL.
  • No centralized API catalog means discovery cost is high.

Evidence and gaps

  • Policy doc: "Level1: readable as Excel/CSV" (Digital Agency PDF) — good requirement, but enforcement and migration paths unclear.
  • Local systems standardization (soumu.go.jp) aims to unify core systems and move to Government Cloud, but transition timelines and operation cost mitigations need measurable SLAs.
  • e-Stat (example portal) shows a mature API approach — compare and contrast with smaller municipalities that only publish PDFs.

Engineering checklist to close the gap

  • API‑first publication: every dataset must have a stable REST/GraphQL endpoint and an OpenAPI spec.
  • Schema registry & JSON Schema: publish machine‑readable schemas and example payloads; use semantic versioning.
  • CI for data quality: automated checks (null rates, column types, unique keys) on every publish.
  • Reference implementations: open‑source client libs (Python/pandas, JS/Fetch) and replication examples.
  • Migration playbook: tools to convert legacy PDFs/Excel to canonical CSV with provenance metadata.
  • Code snippets (quick wins):

    • Fetch CSV and validate with pandas (Python):
    import pandas as pd
    

    from jsonschema import validate

    url = 'https://example.gov/dataset.csv'

    df = pd.read_csv(url)

    schema = {

    "type": "array",

    "items": {"type": "object",

    "properties": {"id": {"type": "integer"},

    "date": {"type": "string", "format": "date"}},

    "required": ["id","date"]}

    }

    Convert to records and validate first 100 rows

    for row in df.head(100).to_dict(orient='records'):

    validate(row, schema)

    • Simple curl to fetch OpenAPI spec:
    curl -s https://example.gov/openapi.json | jq '.paths'

    Policy metrics to track

    • % of datasets available as CSV/JSON vs PDF
    • % datasets with OpenAPI/JSON Schema
    • Median time to publish updates (pipeline latency)
    • Number of CI failures prevented per month

    Implementation roadmap (90 days)

    • Phase 1: Inventory & quick wins — catalog current datasets, convert top 20 PDFs to CSV + schemas.
    • Phase 2: Platform — deploy central schema registry, API gateway, public catalog.
    • Phase 3: Developer ecosystem — publish client SDKs, sandbox, and hackathons for third‑party reuse.

    まとめ

    Policy moves (machine‑readability rules, system standardization) are happening and that's promising. But engineers know policy ≠ product. To make data actually usable, we need API contracts, schema governance, CI on data, and open reference tooling. These are implementable steps that turn the manifesto into production reality.

    おかむーから一言

    I’ve built startups and shipped messy APIs at 2am — trust me, clean contracts + CI make governance scalable. Let’s stop downloading PDFs and start calling APIs!