Code-Driven Manifesto: A Technical Audit of Municipal Open Data in Japan

IT Policy Proposals
Code-Driven Manifesto: A Technical Audit of Municipal Open Data in Japan

どうも〜おかむーです! Today we dig into how Japanese local governments publish data — from PDFs tucked away on websites to dashboards that actually help people. エンジニア的に言うと、データ公開は「API一本で解決する」話なんですよ。

  • Many municipalities publish datasets but often in PDF/locked formats.
  • Digital Agency and Ministry of Internal Affairs (Soumu) provide rules and case studies — but adoption is uneven.
  • Technical fixes (APIs, DCAT, validation pipelines) are low-hanging fruit that boost reuse and accountability.

結論

Japan has the right policy push — see Soumu and Digital Agency resources (e.g. digital.go.jp, soumu.go.jp) — but in practice machine-readability and API-first publication are inconsistent across municipalities. 要するに、公開はしているけど“使える形”で出ていない。これを工程化してCI/CDで回せば、データ活用が一気に伸びますよね。

Report

Background & evidence

Look, the Digital Agency publishes local case studies (https://digital.go.jp/resources/data_case_study_local) and Soumu defines machine-readable criteria (https://www.soumu.go.jp/menu_seisaku/ictseisaku/ictriyou/opendata/). GovTech Tokyo runs practical dashboards (https://www.govtechtokyo.or.jp) — great examples. Internationally, flood-risk apps in the US combine census, NOAA and elevation data to give 95% coverage of residents' flood risk (sorabatake.jp), showing what cross-dataset reuse enables.

Technical gaps observed

  • Formats: PDFs and image-scanned tables remain widespread. The Digital Agency draft rules explicitly call for CSV/Excel as Level 1 machine-readable formats.
  • Metadata: lack of standard metadata (DCAT, schema.org) makes discovery hard.
  • APIs: many datasets have no REST/GraphQL endpoints; bulk CSV is the only option.
  • Quality: inconsistent encodings, merged Excel cells, missing timezones, and undocumented column semantics.

Concrete engineering checks

  • Automated validation: use frictionless (python) or csvlint in CI to assert schema and types.
  • Metadata: publish DCAT JSON and JSON-LD for each dataset for cataloging and SEO.
  • API gateway: wrap CSV endpoints with a simple API (FastAPI) that serves JSON/CSV and supports filtering/pagination.

Example: quick Python snippet to load a CSV and validate with frictionless

import requests

from frictionless import Resource

r = requests.get('https://example.gov/dataset.csv')

open('dataset.csv','wb').write(r.content)

resource = Resource('dataset.csv')

report = resource.validate()

print(report.valid)

Fallback for PDFs: use tabula-py or Camelot to extract tables, but treat this as temporary — automated PDF->CSV converters are brittle.

Policy vs practice: measurable gaps

Policy documents from Soumu/Digital Agency set expectations for machine-readable publication and cross-department reuse, but case studies show adoption is patchy — some cities have polished dashboards (GovTech Tokyo work), others only post PDFs. The metric to track is simple: percentage of datasets in a municipal catalog that are machine-readable CSV/JSON + have API endpoints and DCAT metadata.

Roadmap / Recommendations

  • Inventory: run a crawler to classify dataset formats and produce a per-municipality "machine-readability" score.
  • Pipeline: implement GitOps for data (CSV files in a repo, validations triggered on PRs, automatic publish to portal).
  • Schema catalog: define canonical schemas for common domains (facilities, budgets, disasters) and publish JSON Schema + example data.
  • Incentives: central dashboards that surface reuse metrics (downloads, API calls) to reward municipalities.

まとめ

  • Japan has policy foundations and good case studies, but machine-readability and API coverage need scaling.
  • Technical fixes (validation, API wrappers, metadata standards) are straightforward and high-impact.
  • Measure progress with automated crawlers and a clear per-municipality score — then iterate.

おかむーから一言

データは出すだけじゃ意味がないんですよ。エンジニア目線でパイプライン化して、誰でも使えるインタフェースにするのが肝です!