Code Speaks: How Prefectural Data Should Behave (Nagano as a Case Study)

IT Policy Proposals
Code Speaks: How Prefectural Data Should Behave (Nagano as a Case Study)

どうも〜おかむーです! Hey — okamu here! Today I want to dig into how prefectures and city governments publish data and systems, and what we can do as engineers and civic tech folks to make them actually useful.

  • Many local pages (e.g., pref.nagano.lg.jp) publish lots of policy documents but often as PDFs, not machine-readable data.
  • Gov bodies like Digital Agency and GovTech Tokyo are pushing dashboards and shared tooling, but operational gaps remain (APIs, consistency, provenance).
  • Engineering fixes (APIs, CSV/JSON first, OpenAPI, CI for data) can close KPI monitoring gaps and unlock reuse.

結論

Public sector digital services should make their policy numbers and KPIs available as first-class, machine-readable artifacts. Engineer-wise, this means CSV/JSON + stable REST APIs with schema and provenance, automated ETL pipelines to power dashboards, and open specs (OpenAPI / DCAT) so external developers can build on top. Nagano and many municipalities are getting there, but PDFs and ad-hoc tables are still the norm, which wastes time and hides accountability.

Report: what I looked at and what it implies

Sources and context

I checked the Nagano Prefectural Government site (https://www.pref.nagano.lg.jp/) as an example of a prefectural portal; Digital Agency (https://www.digital.go.jp/) and GovTech Tokyo resources (https://www.govtechtokyo.or.jp/) as national and municipal-level programs that push data reuse. Also lean on e-Stat (Japan's statistical API) as the canonical machine-readable national dataset provider.

Engineering note: if a policy references population, budget, or KPI numbers, treat those numbers like code — they must be versioned, testable, and machine-readable.

Problem patterns I keep finding

  • PDF-first publishing: policy reports, KPI tables, and tender documents are often PDF scans. Check this out — a CSV would let you compare year-to-year in 30 seconds; PDFs force manual extraction. In short: PDFs = friction.
  • No or limited APIs: some cities publish dashboards but not APIs. Dashboards are great for humans, but engineers need endpoints to automate monitoring and validation.
  • Inconsistent identifiers: datasets lack stable keys (e.g., JIS codes for municipalities), so joins across datasets become fragile.
  • Missing metadata: provenance, update timestamps, and license info are frequently absent or buried.

Example technical checks (what I'd run as an engineer)

  • Check landing pages for an OpenData catalog (DCAT JSON) or links to CSV/JSON files under a /data/ path.
  • Confirm CORS and API rate limits on any dashboard endpoints.
  • Try the e-Stat API for demographic baselines: e-Stat provides restful endpoints (https://www.e-stat.go.jp/en/api) — using them as canonical reference simplifies reconciliation.

Code snippet: fetch a CSV from an open dataset and inspect headers (Python)

import requests

import pandas as pd

url = 'https://example.prefecture.gov/data/population.csv'

r = requests.get(url)

r.raise_for_status()

df = pd.read_csv(pd.compat.StringIO(r.text))

print(df.head())

If the site only has a PDF, use tabula or Camelot to extract tables programmatically, but that's brittle. Example: using tabula-py to extract the first table

import tabula

tables = tabula.read_pdf('policy_report.pdf', pages=1, multiple_tables=True)

len(tables), tables[0].head()

KPI and policy target gaps

A common policy flow: central government (Digital Agency) sets funding and program goals (e.g., digital infrastructure KPIs), prefectures apply and publish outcomes. But often the published "実績 (results)" are in narrative reports or PDFs, not machine-readable time-series. That makes it hard to compute gaps between target and actual program outcomes automatically.

Engineer-wise, you want:

  • targets.csv: columns {program_id, year, target_value}
  • results.csv: columns {program_id, year, measured_value, source_url}
  • a reconciler script that computes gaps and alerts when slippage exceeds thresholds

This is trivial to implement with a scheduled CI job that runs tests on incoming data and updates a dashboard.

Quality of formats and interoperability

  • Use JSON-LD or schema.org for metadata so search engines and aggregators can index datasets.
  • Publish an OpenAPI spec for any REST endpoints. That helps clients generate type-safe clients and test suites.
  • Adopt DCAT for dataset catalogs so GovTech Tokyo-style shared dashboards can automatically harvest datasets.

Privacy and security

When publishing microdata, ensure proper anonymization and differential privacy where needed. For APIs, require API keys and rate limiting for heavy endpoints; but non-sensitive aggregates should be fully public with permissive licenses (CC-BY or similar).

Operational suggestions (practical roadmap)

  • Inventory: generate a machine-readable inventory (DCAT) of every dataset referenced in policy docs. Tools like CKAN or data.gouv.jp can help.
  • Lift PDFs to CSV/JSON: prioritize KPI tables, budget tables, and regulatory lists. Automate extraction and human-verify once.
  • Expose APIs with OpenAPI and examples. Use schema validation (JSON Schema) on ingest.
  • CI/CD for data: run tests (schema, ranges, time continuity) and publish versioned snapshots (S3 with signed URLs or a data registry).
  • Dashboard from canonical APIs, not manual exports. Dashboard UI is for humans; API is for machines.
  • Code snippet: simple gap checker (pseudocode)

    # assumes targets.csv and results.csv already published
    

    import pandas as pd

    targets = pd.read_csv('targets.csv')

    results = pd.read_csv('results.csv')

    merged = targets.merge(results, on=['program_id','year'])

    merged['gap_pct'] = (merged['measured_value'] - merged['target_value']) / merged['target_value']

    alerts = merged[merged['gap_pct'] < -0.1] # >10% shortfall

    alerts.to_csv('alerts.csv')

    Reuse and civic innovation

    Open, well-documented machine-readable data enables:

    • automated watchdogs that detect KPI slippage
    • civic apps that combine datasets (transport, demographics, budget) using stable IDs
    • researchers to produce reproducible analyses without manual scraping

    GovTech Tokyo's efforts to centralize dashboards are a great step, but they need to be backed by API-first publishing from each prefecture/city.

    まとめ

    Check this out: the technical barriers to making government data truly useful are small and well-understood. Move from PDF-first to data-first, provide APIs with schema and provenance, and run CI on datasets. That turns policy numbers from static text into living, testable artifacts we can all rely on.

    おかむーから一言

    I've built products from zero and shipped GovTech projects — tech can make government measurable and improvable. Let's treat policy numbers like code: testable, versioned, and automatable. If you're in a city hall reading this, start with your next KPI table as CSV and watch the magic happen!