Code-Driven Manifesto: Evaluating Japan's Local Open Data & Digital Grants

IT Policy Proposals
Code-Driven Manifesto: Evaluating Japan's Local Open Data & Digital Grants

どうも〜おかむーです! Hey — okamu here! Today we're doing a bit of engineering-style audit of Japanese local open data and theDigital Denen Toshi funding ecosystem, looking at catalogs, CSV quality, APIs and whether policy numbers actually map to machine-readable reality.

  • Cities publish datasets but often as ad-hoc CSV/XLSX or PDFs, hurting reuse
  • API coverage and standardized metadata (DCAT/Data Package) are uneven across prefectures
  • Simple technical fixes (UTF-8, schema, APIs) unlock measurable policy transparency

結論

Local governments publish lots of useful data (Tokyo catalog, Saitama portal, Niigata CSV guidance, Digital Agency programs), but friction points—file formats, encodings, missing metadata, lack of APIs and KPI time-series—block civic reuse. 要するに、データを"open"にするなら、人間向けPDFじゃなくてエンジニア向けのAPIとSchemaを出せば状況が激変します。

Report

What I looked at

  • Tokyo Open Data Catalog (catalog.data.metro.tokyo.lg.jp): many CSV/XLSX datasets but varying metadata completeness
  • Saitama Open Data Portal: catalog + downloadable files but mixed formats
  • Niigata CSV manual: explicit guidance — good signal!
  • Digital Denen Toshi grants pages and Sukagawa evaluation report: show funding -> outcomes, but data is fragmented

Common technical problems (これ見てくださいよ)

  • PDF-embedded tables or XLSX only: not machine-first
  • Encodings: some CSVs in Shift_JIS or unspecified → parsing errors
  • Missing machine-readable metadata: no DCAT, no license field, no update frequency
  • No stable APIs or rate-limited ad-hoc CSV downloads → hard to pipeline
  • Schema drift: inconsistent column names, date formats, mixed Japanese era and Gregorian

要するに、データの実務で一番困るのは"不安定なスキーマ"と"不明瞭なライセンス"ということです。

Quick technical checks & reproducible steps

Example: fetch CSV, normalize encoding and parse with pandas (Python). Code sample:

import pandas as pd

from io import BytesIO

import requests

r = requests.get('https://example.pref.jp/dataset.csv')

ensure UTF-8

df = pd.read_csv(BytesIO(r.content), encoding='shift_jis', parse_dates=['date_col'])

normalize column names

df.columns = [c.strip().lower().replace(' ', '_') for c in df.columns]

print(df.head())

要するに、エンジニア的に言うとfetch->normalize->validateが基本です。

Policy vs data: KPI gaps

Digital Denen Toshi grants and local strategies state targets (e.g. digital implementation, tourism uplift), but publicly available evaluation (e.g. Sukagawa) often gives qualitative reports or PDFs. Without time-series, per-project KPIs in machine-readable form, automated monitoring is impossible. Suggest: publish per-grant metrics as JSON/CSV with schema: project_id, metric_name, target_value, measured_value, date.

Concrete improvements (engineering roadmap)

  • Adopt DCAT-AP-JP / Data Packages for catalog metadata (license, update frequency, schema)
  • Prefer UTF-8-without-BOM CSV or JSON/NDJSON; avoid PDFs for raw data
  • Provide RESTful APIs + OpenAPI spec and a bulk-download S3 endpoint (CDN)
  • Use semantic column names, ISO dates, and stable IDs (UUIDs) for records
  • Publish evaluation KPIs as time-series CSV/JSON and expose dashboards
  • Provide example clients (Python/R/JS) and rate-limited API keys for reproducible civic apps

まとめ

Japan's municipalities have the raw material — catalogs, guides (Niigata), and funding streams — but need standardized metadata, consistent encodings, APIs and machine-readable KPI disclosures to turn policy into verifiable outcomes. Small engineering investments (Data Packages, OpenAPI, UTF-8, stable schema) massively increase civic reuse and policy accountability.

おかむーから一言

I've built products and startups; trust me, making data machine-friendly is the lowest-friction, highest-leverage move governments can make. Let's ship APIs, not PDFs!

Code-Driven Manifesto: Fixing Japan's Open Data Gaps with APIs and Schemas | Japan Political Economy Global Report