Code Speaks: Auditing Japanese Local Gov Data — CSVs, APIs and Practical Fixes

IT Policy Proposals
Code Speaks: Auditing Japanese Local Gov Data — CSVs, APIs and Practical Fixes

どうも〜おかむーです! Hi everyone — today I want to take a tech-first look at how Japanese local governments publish data, what trips developers up, and concrete engineering fixes that actually work.

  • Many prefectures and cities publish datasets (CSV catalogs exist) but formats and metadata are inconsistent
  • PDFs and unstructured releases still block automated analysis; APIs are unevenly available
  • Practical path: canonical schema + CSVW/JSON-LD + lightweight APIs + CI for data quality

結論

Local governments often publish valuable datasets (see Tokyo catalog, Niigata, Saitama, Hakodate), but the translation from "published" to "machine-consumable and auditable" is incomplete. 要するに、データはあるけどエンジニアが使うにはつらい。短期的にできる改善はCSV/encoding normalization、schema validation、OpenAPI endpointsの追加。中長期的にはDCAT/CSVWとバージョン付きAPIを整備すること!

Report: what's actually out there

What's available (good examples)

  • Tokyo Open Data Catalog: https://catalog.data.metro.tokyo.lg.jp/dataset — lots of CSVs and policy-related datasets that can be used for dashboards
  • GovTech Tokyo services: https://www.govtechtokyo.or.jp/services/data-utilization/ — organizational effort to share dashboards and build common visualization
  • Prefectural/city portals: Niigata (CSV manual & datasets: https://www.city.niigata.lg.jp/shisei/seisaku/it/open-data/index.html and https://www.city.niigata.lg.jp/shisei/seisaku/it/open-data/index.files/csv_manual_v1.1.pdf), Saitama datasets (https://opendata.pref.saitama.lg.jp/datasets), Hakodate CSV listing (https://www.harp.lg.jp/opendata/dataset/79.html)

These sindicates show intent and capacity. But intent ≠ usable data.

Common technical problems (これ見てくださいよ)

  • PDF-first reporting: many policy targets and performance reports are delivered as PDF, sometimes with tables embedded. That blocks automation.
  • Encoding & format drift: CSVs mix encodings (Shift_JIS vs UTF-8), inconsistent column names, date formats (YYYY/MM/DD vs YYYY-MM-DD vs Japanese era), and missing headers.
  • Missing machine-readable metadata: no schema (types, units), no license tags, no update timestamps.
  • Sparse APIs: some dashboards exist, but few provide stable, versioned REST/GraphQL APIs with OpenAPI specs.
  • Domain & trust issues: site structures vary (some countries use .gov/.gov.cn or national vs local domains), affecting centralized discovery and crawler reliability (see discussion about .gov vs .gov.cn on public forums).

Why this matters (policy evaluation angle)

Policy documents often state numeric targets (e.g., participation rates, budgetary ceilings, disaster preparedness goals), but if the underlying operational datasets lack time stamps, consistent identifiers, or are locked in PDFs, it's hard to compute progress or replicate the government's claims. 要するに、「目標はあるけど検証できない」状態になっているんですよね。

Technical checks and quick recipes

1) Automated ingestion checklist

  • Detect encoding and normalize to UTF-8
  • Parse dates robustly with dateutil or pandas.to_datetime
  • Validate schema using a library (pandera, jsonschema or CSVW descriptions)
  • Persist raw and cleaned versions with provenance (commit hashes or dataset versions)

Here's a minimal Python snippet showing robust CSV ingestion and schema validation with pandas + pandera:

import chardet

import pandas as pd

import pandera as pa

from io import BytesIO

import requests

url = 'https://example.local.gov/dataset.csv'

raw = requests.get(url).content

enc = chardet.detect(raw)['encoding'] or 'utf-8'

df = pd.read_csv(BytesIO(raw), encoding=enc)

schema = pa.DataFrameSchema({

'date': pa.Column(pa.DateTime, coerce=True),

'region_id': pa.Column(pa.Int),

'value': pa.Column(pa.Float)

})

df = schema.validate(df)

print(df.head())

要するに、最初にエンコーディングと日付を正規化しておくとあとの処理が捗ります。

2) From CSV to API in 30 minutes

Use FastAPI to wrap cleaned CSV into a simple JSON endpoint and publish OpenAPI automatically.

from fastapi import FastAPI

import pandas as pd

app = FastAPI()

df = pd.read_csv('cleaned.csv', parse_dates=['date'])

@app.get('/api/v1/data')

def get_data(start: str = None, end: str = None):

q = df

if start:

q = q[q['date'] >= start]

if end:

q = q[q['date'] <= end]

return q.to_dict(orient='records')

This gives immediate value to internal analysts and external civic tech teams.

3) Comparing policy targets vs outcomes (method)

  • Extract targets from policy PDFs. If structured text is not available, use OCR+table detection (tabula-py or Camelot) to extract numeric targets
  • Align target identifiers with dataset keys (e.g., region code, fiscal year)
  • Compute gap = (target - actual) / target and visualize with a simple charting library or dashboard (e.g., Grafana or Superset)

Pseudo-workflow:

  • OCR PDF for target table -> CSV
  • Normalize keys and dates
  • Join with operational dataset
  • Aggregate and compute gaps
  • Policy & governance suggestions (concrete)

    Short term (weeks):

    • Provide CSV downloads with UTF-8, include header rows with clear English/Japanese column names and units
    • Add dataset-level JSON metadata: title, description, contact, license, last_updated
    • Publish a small OpenAPI wrapper for top 5 datasets

    Medium term (3-12 months):

    • Adopt DCAT + CSVW/JSON-LD for catalog metadata so catalogs are machine-discoverable
    • Version datasets and keep raw snapshots in object storage (S3/GCS) with immutable URIs
    • Implement a CI job that validates schema and alerts when a dataset changes unexpectedly

    Long term (1-2 years):

    • Centralized discovery with federated search across lg.jp/go.jp domains and clear domain policy (.lg.jp for local gov, .go.jp for central?)
    • Stable, authenticated APIs for administrative data with rate limits and usage metrics

    Tools and standards to adopt

    • CSVW (CSV on the Web) for schema and metadata
    • DCAT for dataset cataloging
    • OpenAPI for API endpoints
    • Pandera/jsonschema for validation
    • CKAN/DataHub if you want a packaged catalog solution

    まとめ

    • There is a lot of public data: Tokyo, Niigata, Saitama, Hakodate examples show good starts
    • Biggest blockers: PDFs, inconsistent CSVs (encoding, date formats), missing metadata and APIs
    • Engineering fixes are straightforward: normalize encodings, validate schemas, add lightweight APIs, and automate tests

    おかむーから一言

    テクノロジーで社会をアップデートするって、結局は「小さな改善を積み重ねて誰でも再現できる状態」にすることなんです。エンジニア的に言うと、データをCIで守る文化を入れましょう!