Code-speaking Manifesto: Evaluating Japan's Public Data Ecosystem

IT Policy Proposals
Code-speaking Manifesto: Evaluating Japan's Public Data Ecosystem

どうも〜おかむーです! Today I'm taking a technical scalpel to how Japanese government and local authorities publish data — looking at formats, APIs, KPIs and what engineers actually need to ship useful civic apps.

  • Government often publishes factsheets as PDF but provides occasional CSVs or portals like RAIDA and Digital Agency assets
  • Machine-readability, APIs, and consistent schemas are the main blockers for reusing public data
  • Practical steps: canonical CSV/JSON endpoints, encoding hygiene, schema validation, KPI publishing and CI for datasets

結論

Public agencies are moving toward machine-readable publishing (see RAIDA, Digital Agency CSVs, notice.go.jp), but inconsistency in formats, encodings, and absence of well-documented APIs make real reuse fragile. 要するに、エンジニア的に言うと“one canonical API + schema + CI”が不足しているということです!

Technical report

What I looked at (concrete sources)

  • notice.go.jp CSV: https://notice.go.jp/docs/status_notice.csv — a public CSV endpoint you can curl
  • Digital Agency assets (hospital CSV example): https://www.digital.go.jp/assets/.../xxxxxx_hospital.csv — example of CSV asset hosting
  • RAIDA: https://raida.go.jp/ — evaluation platform for regional digital grants
  • Cabinet Office CSVs: https://www5.cao.go.jp/j-j/wp/wp-je24/csv/*.csv — classic example where both PDF and CSV exist

These show a positive trend: CSVs are available alongside PDFs. But availability ≠ usability.

Key technical issues

  • Encoding and delimiters: many CSVs from gov sites use SHIFT_JIS or include BOMs. If you curl + pandas without handling encoding you get mojibake. 要するに、encoding matters.
  • Schema drift and undocumented fields: column names change between releases, units are not always explicit (counts vs. thousands), timestamps vary.
  • Lack of APIs and pagination: static CSV dumps are fine for snapshot analysis, but real-time apps need REST/GraphQL endpoints with stable URLs and pagination.
  • Metadata and provenance: no DCAT/JSON-LD manifest in many catalogs — hard to automate discovery and versioning.

Quick reproducible checks (engineer notes)

Curl a CSV and inspect encoding:

curl -sS https://notice.go.jp/docs/status_notice.csv -o status_notice.csv

file status_notice.csv

iconv -f SHIFT_JIS -t UTF-8 status_notice.csv > status_notice.utf8.csv

Python snippet to load robustly:

import pandas as pd

from io import BytesIO

import requests

r = requests.get('https://notice.go.jp/docs/status_notice.csv')

guess encoding then decode

text = r.content.decode('cp932', errors='replace')

df = pd.read_csv(BytesIO(text.encode('utf-8')))

print(df.head())

KPI and policy alignment

Digital田園都市交付金 (digital implementation grants) require KPI reporting (see Ministry docs). The problems I found:

  • KPIs often published in PDFs or embedded tables rather than machine-readable time series
  • No standardized KPI schema across municipalities, making national aggregation painful
  • Missing machine-friendly links between grant project records (RAIDA) and periodic performance datasets

Result: policy evaluation becomes manual, slow, and error-prone.

Improvements engineers want

  • Canonical API per dataset with OpenAPI spec, CORS, and stable versioned endpoints
  • CSV/JSON endpoints with UTF-8 and clear column metadata (units, types, enums)
  • DCAT-compliant dataset manifests and dataset versioning (e.g., /catalog/{id}/versions)
  • CI for data pipelines: automated validation tests that run on ingest (schema checks, row counts, null thresholds)
  • Publish KPI time-series as JSON Table Schema or CSVW so aggregators can roll up nationally

Concrete proposal (minimal implementation plan)

  • Add a dataset manifest (JSON-LD/DCAT) for each project page (Digital Agency, RAIDA entries).
  • Serve canonical JSON endpoints in addition to CSV; provide OpenAPI and examples.
  • Enforce UTF-8 output and include CSVW/JSON Table Schema files alongside CSVs.
  • Run nightly CI checks: encoding, schema conformance, row-count anomalies; publish the CI badge/status.
  • Code sketch for a validation test (pseudocode):

    # run as CI job
    

    schema = load_table_schema('schema.json')

    df = pd.read_csv('data.csv', encoding='utf-8')

    assert set(schema['fields']) <= set(df.columns)

    assert df['date'].is_monotonic_increasing()

    assert df['value'].notnull().mean() > 0.95

    Open-data use cases unlocked

    • Cross-municipal KPI dashboards (automated) combining RAIDA project metadata with time-series performance
    • Real-time alerting when grant outcomes deviate from targets (requires APIs and stable schemas)
    • Civic apps that blend hospital locations (digital.go.jp CSV) with local service availability

    まとめ

    • Japan's public sector is publishing more machine-readable artifacts — good!
    • But inconsistent encodings, lack of API contracts, and missing metadata block scale reuse
    • Practical, low-cost fixes (UTF-8, JSON endpoints, schema files, CI checks) would unlock a lot of civic innovation

    おかむーから一言

    これ、やるかやらないかは行政の“エンジニア感”の差です。コードで約束を守れば、政策評価はもっと速く、透明になりますよ!