Code Speaks: Auditing Japanese Local Open Data from an Engineer’s Lens

IT Policy Proposals
Code Speaks: Auditing Japanese Local Open Data from an Engineer’s Lens

どうも〜おかむーです! Today I want to take a quick, engineer-first look at how Japanese national and local governments publish public data. We're doing this under the banner “Code Speaks Manifesto” — meaning: treat policy as code, validate it with data, and fix the plumbing if it breaks!

  • This inspects open data practices across Tokyo, Saitama, Niigata, Chūō-ku and central gov sources
  • Main problems: mixed formats (PDFs), inconsistent CSVs, sparse APIs and weak metadata
  • Practical fixes: API-first, CSVW/JSON-LD metadata, validation pipelines and KPI dashboards

結論

Public datasets exist and can be incredibly useful, but many Japanese local/national portals still ship data like documents, not products. From an engineering perspective, the win is clear: make datasets machine-first (CSV/JSON + schema + API), publish metadata (DCAT/CSVW), and add automated validation so policy KPI tracking becomes reliable and auditable.

Report

What I looked at

These are representative sources found on public portals:

  • Tokyo Open Data Catalog (catalog.data.metro.tokyo.lg.jp) — CSVs such as a disaster-awareness survey
  • Saitama Open Data portal (opendata.pref.saitama.lg.jp) — catalog-driven downloads
  • Niigata CSV guidance (city.niigata.lg.jp CSV manual) — guidance for machine-friendly CSV
  • Chūō-ku open data page (city.chuo.lg.jp) — mentions CSV-first publishing
  • Central government/stake reports (digital.go.jp, chisou.go.jp guidance, notice.go.jp CSV)

These show a positive trend: availability of CSV and catalogs. But UX and tooling vary widely.

Common technical issues

  • PDF-first disclosure: important reports end up as PDFs, not machine-readable tables. 要するに、データが取り出しにくいということです。
  • Inconsistent CSV schemas: column names, encodings (Shift_JIS vs UTF-8), missing metadata (units, timestamps).
  • Weak metadata: catalogs often lack machine-readable schema (no CSVW/JSON-LD), so programmatic discovery is hard.
  • Limited APIs: many portals only allow file downloads. No REST endpoints or queryable stores for filtered time-series.
  • KPI traceability gap: policy docs (e.g., digital田園都市 guidance) list KPIs but linking KPIs to published datasets is often manual.

Engineer-friendly fixes (concrete)

  • Publish machine-readable metadata
  • - Use DCAT for dataset discovery and CSVW/JSON-LD for column-level schema.

  • Provide APIs and data slices
  • - Add simple endpoints (e.g., /api/v1/datasets/{id}/rows?from=2023-01-01) or adopt CKAN/SODA.

  • Standardize encoding & schema
  • - UTF-8 default, ISO 8601 timestamps, explicit units and controlled vocabularies.

  • Automated validation pipeline
  • - CI job that validates CSVs with goodtables, csv-validator, and publishes a freshness/quality badge.

  • KPI dashboards that link policy -> dataset -> transformation
  • - Link each KPI to the dataset and the transformation SQL/notebook that computes it.

    Example: quick Python fetch of a CSV (engineer-style):

    import requests
    

    import pandas as pd

    url = 'https://catalog.data.metro.tokyo.lg.jp/dataset/example.csv'

    r = requests.get(url)

    r.encoding = 'utf-8'

    df = pd.read_csv(pd.compat.StringIO(r.text))

    print(df.head())

    This is the sort of one-liner that fails when encoding or schema is inconsistent. So automate the checks!

    Policy vs data: closing the loop

    Policy docs (e.g., cross-government KPI guidance) demand measurable outcomes. If datasets are buried in PDFs or inconsistent CSVs, you cannot reliably measure progress. The easy engineering rule: every KPI must have (a) a canonical dataset id, (b) a documented query, and (c) an automated freshness check.

    まとめ

    These portals are doing the right thing by publishing data, but to make data actionable we need product-level improvements: machine-readable metadata, stable APIs, schema validation, and explicit KPI-to-dataset mapping. Do that and policy becomes auditable, computable, and improvable.

    おかむーから一言

    Tech can make government measurable and improvable — that's why I build. Let's push for APIs, schemas, and automated tests so policy can actually be tested like code!