Code Speaks: How Prefectural Data Should Behave (Nagano as a Case Study)

どうも〜おかむーです! Hey — okamu here! Today I want to dig into how prefectures and city governments publish data and systems, and what we can do as engineers and civic tech folks to make them actually useful.
- Many local pages (e.g., pref.nagano.lg.jp) publish lots of policy documents but often as PDFs, not machine-readable data.
- Gov bodies like Digital Agency and GovTech Tokyo are pushing dashboards and shared tooling, but operational gaps remain (APIs, consistency, provenance).
- Engineering fixes (APIs, CSV/JSON first, OpenAPI, CI for data) can close KPI monitoring gaps and unlock reuse.
結論
Public sector digital services should make their policy numbers and KPIs available as first-class, machine-readable artifacts. Engineer-wise, this means CSV/JSON + stable REST APIs with schema and provenance, automated ETL pipelines to power dashboards, and open specs (OpenAPI / DCAT) so external developers can build on top. Nagano and many municipalities are getting there, but PDFs and ad-hoc tables are still the norm, which wastes time and hides accountability.
Report: what I looked at and what it implies
Sources and context
I checked the Nagano Prefectural Government site (https://www.pref.nagano.lg.jp/) as an example of a prefectural portal; Digital Agency (https://www.digital.go.jp/) and GovTech Tokyo resources (https://www.govtechtokyo.or.jp/) as national and municipal-level programs that push data reuse. Also lean on e-Stat (Japan's statistical API) as the canonical machine-readable national dataset provider.
Engineering note: if a policy references population, budget, or KPI numbers, treat those numbers like code — they must be versioned, testable, and machine-readable.
Problem patterns I keep finding
- PDF-first publishing: policy reports, KPI tables, and tender documents are often PDF scans. Check this out — a CSV would let you compare year-to-year in 30 seconds; PDFs force manual extraction. In short: PDFs = friction.
- No or limited APIs: some cities publish dashboards but not APIs. Dashboards are great for humans, but engineers need endpoints to automate monitoring and validation.
- Inconsistent identifiers: datasets lack stable keys (e.g., JIS codes for municipalities), so joins across datasets become fragile.
- Missing metadata: provenance, update timestamps, and license info are frequently absent or buried.
Example technical checks (what I'd run as an engineer)
- Check landing pages for an OpenData catalog (DCAT JSON) or links to CSV/JSON files under a /data/ path.
- Confirm CORS and API rate limits on any dashboard endpoints.
- Try the e-Stat API for demographic baselines: e-Stat provides restful endpoints (https://www.e-stat.go.jp/en/api) — using them as canonical reference simplifies reconciliation.
Code snippet: fetch a CSV from an open dataset and inspect headers (Python)
import requests
import pandas as pd
url = 'https://example.prefecture.gov/data/population.csv'
r = requests.get(url)
r.raise_for_status()
df = pd.read_csv(pd.compat.StringIO(r.text))
print(df.head())
If the site only has a PDF, use tabula or Camelot to extract tables programmatically, but that's brittle. Example: using tabula-py to extract the first table
import tabula
tables = tabula.read_pdf('policy_report.pdf', pages=1, multiple_tables=True)
len(tables), tables[0].head()
KPI and policy target gaps
A common policy flow: central government (Digital Agency) sets funding and program goals (e.g., digital infrastructure KPIs), prefectures apply and publish outcomes. But often the published "実績 (results)" are in narrative reports or PDFs, not machine-readable time-series. That makes it hard to compute gaps between target and actual program outcomes automatically.
Engineer-wise, you want:
- targets.csv: columns {program_id, year, target_value}
- results.csv: columns {program_id, year, measured_value, source_url}
- a reconciler script that computes gaps and alerts when slippage exceeds thresholds
This is trivial to implement with a scheduled CI job that runs tests on incoming data and updates a dashboard.
Quality of formats and interoperability
- Use JSON-LD or schema.org for metadata so search engines and aggregators can index datasets.
- Publish an OpenAPI spec for any REST endpoints. That helps clients generate type-safe clients and test suites.
- Adopt DCAT for dataset catalogs so GovTech Tokyo-style shared dashboards can automatically harvest datasets.
Privacy and security
When publishing microdata, ensure proper anonymization and differential privacy where needed. For APIs, require API keys and rate limiting for heavy endpoints; but non-sensitive aggregates should be fully public with permissive licenses (CC-BY or similar).
Operational suggestions (practical roadmap)
Code snippet: simple gap checker (pseudocode)
# assumes targets.csv and results.csv already published
import pandas as pd
targets = pd.read_csv('targets.csv')
results = pd.read_csv('results.csv')
merged = targets.merge(results, on=['program_id','year'])
merged['gap_pct'] = (merged['measured_value'] - merged['target_value']) / merged['target_value']
alerts = merged[merged['gap_pct'] < -0.1] # >10% shortfall
alerts.to_csv('alerts.csv')
Reuse and civic innovation
Open, well-documented machine-readable data enables:
- automated watchdogs that detect KPI slippage
- civic apps that combine datasets (transport, demographics, budget) using stable IDs
- researchers to produce reproducible analyses without manual scraping
GovTech Tokyo's efforts to centralize dashboards are a great step, but they need to be backed by API-first publishing from each prefecture/city.
まとめ
Check this out: the technical barriers to making government data truly useful are small and well-understood. Move from PDF-first to data-first, provide APIs with schema and provenance, and run CI on datasets. That turns policy numbers from static text into living, testable artifacts we can all rely on.
おかむーから一言
I've built products from zero and shipped GovTech projects — tech can make government measurable and improvable. Let's treat policy numbers like code: testable, versioned, and automatable. If you're in a city hall reading this, start with your next KPI table as CSV and watch the magic happen!
Sources
- https://www.pref.nagano.lg.jp/
- https://metidx-gov.note.jp/n/n9468573c213b
- https://ja.wikipedia.org/wiki/%E6%97%A5%E6%9C%AC%E3%81%AE%E8%A1%8C%E6%94%BF%E6%A9%9F%E9%96%A2
- https://picks-design.com/blog/5751/
- https://www.gyosei.ainosato.works/entry/gyosei-definition
- https://www.zhihu.com/question/40553450
- https://www.govtechtokyo.or.jp/services/data-utilization/
- https://bus.gov.ru/
- https://note.govtechtokyo.jp/n/n77785a8254d6
- https://www.zhihu.com/tardis/zm/art/1924492115896960699
- https://ja.wikipedia.org/wiki/%E3%83%87%E3%82%B8%E3%82%BF%E3%83%AB
- https://www.chisou.go.jp/sousei/pdf/r5_guideline-checkaction.pdf
- https://www.digital.go.jp/
- https://www.city.sukagawa.fukushima.jp/shisei/gyoseiunei/keikaku/chiho_sosei/1015604/4045.html
- https://biz.kddi.com/content/column/smartwork/what-is-digital/
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.