Code Speaks: Auditing Japan’s Public CSVs and Data Practice

どうも〜おかむーです! Today I'm diving into some public CSVs from Japanese government sites with an engineer's eye — "code talks" style. This is a technical take on machine-readability, APIs, and how policy claims map to data.
- This article inspects a few live CSVs (new-car sales, public investment, hospital list, NICTER notices).
- Main finding: data is published, but machine-friendliness, metadata, and APIs are inconsistent — lots of low-hanging UX fixes.
- I give concrete parsing tips, a short Python recipe, and actionable improvements (OpenAPI, CSVW/JSON-LD, catalogs).
Conclusion
Public bodies are publishing useful CSVs (see https://www5.cao.go.jp/j-j/wp/wp-je24/csv/d1-1-4.csv and https://www5.cao.go.jp/j-j/wp/wp-je24/csv/d1-1-42.csv, plus https://www.digital.go.jp/assets/.../xxxxxx_hospital.csv and https://notice.go.jp/docs/status_nicter.csv). But from a developer standpoint, the work stops at dumping files: missing stable APIs, ambiguous encodings/units, absent schema/metadata, and no discoverable machine contracts. 要するに、データはあるけど“API-first”の体験には程遠いということです。
Report
Data sources I eyeballed
- New car sales timeseries (cabinet office): https://www5.cao.go.jp/j-j/wp/wp-je24/csv/d1-1-4.csv
- Public investment budget vs settlement: https://www5.cao.go.jp/j-j/wp/wp-je24/csv/d1-1-42.csv
- Hospital registry CSV (Digital Agency asset): https://www.digital.go.jp/.../xxxxxx_hospital.csv
- NICTER notice feed: https://notice.go.jp/docs/status_nicter.csv
Look at this: some CSVs contain a BOM and Japanese headers — fine for humans, but code needs to know encoding and units!
Technical assessment (quick)
- Formats: Mostly CSV — great for simplicity, but CSV alone lacks typed schema and units. PDF-to-CSV conversions sometimes leave artifacts.
- Encoding: BOM/Shift_JIS/UTF-8 ambiguity observed in snippets. エンジニア的に言うと, you will hit decode errors unless you handle encodings explicitly.
- Metadata: No machine-readable schema (CSVW/JSON-LD) attached. No clear unit fields (e.g., is sales in thousands or ten-thousands?).
- API: No documented REST/GraphQL API for these datasets. That makes live apps fragile (screen-scrape + CSV polling).
- Provenance & licensing: Often implicit; needs explicit license & update cadence fields.
Quick reproducible parsing recipe (Python)
import requests
import io
import pandas as pd
url = 'https://www5.cao.go.jp/j-j/wp/wp-je24/csv/d1-1-4.csv'
res = requests.get(url)
Try utf-8-sig to handle BOM; fall back to shift_jis if needed
data = io.StringIO(res.content.decode('utf-8-sig'))
df = pd.read_csv(data)
print(df.head())
parsing tips: rename columns to ascii keys, add units column based on metadata lookup
Notes: handle utf-8-sig, try shift_jis if decode fails; coerce numeric with pd.to_numeric(errors='coerce').
Policy-to-data mapping: how to check targets vs reality
This workflow lets you verify claims like “public investment increased X%” against machine data.
Concrete improvement proposals
- Publish OpenAPI-backed JSON endpoints for each dataset, with stable versioning.
- Attach CSVW metadata (tableSchema) or JSON Schema + example payloads. Include: units, frequency, last_updated, contact, license.
- Standardize encoding to UTF-8 (no BOM) and UTF-8 Content-Type headers.
- Provide dataset catalog (CKAN or DataPortal) with machine-searchable tags; link to Japan Dashboard/e-Stat.
- Ship example ETL scripts (Python/R) in a public GitHub repo so civictech teams can contribute.
- Offer webhooks or push-notifications for update events (or simple ETag/Last-Modified headers).
まとめ
- Good: government publishes raw CSVs — that's the baseline win!
- Bad: missing metadata, inconsistent encodings, and no API contract make rebuilding dashboards fragile.
- Actionable: add schema, standardize UTF-8, publish OpenAPI/CSVW, and provide example code — these are high-ROI fixes.
おかむーから一言
As an entrepreneur-engineer, I want gov data to be as easy to import as pip install. Let's make data discoverable, verifiable, and API-first — code should be the manifest of policy, not a guessing game.
Sources
- https://www5.cao.go.jp/j-j/wp/wp-je24/csv/d1-1-4.csv
- https://notice.go.jp/docs/status_nicter.csv
- https://www.digital.go.jp/assets/contents/node/basic_page/field_ref_resources/0066e8a8-6734-44ab-a9a9-8e09ba9cb508/xxxxxx_hospital.csv
- https://www.jinji.go.jp/content/900024615.csv
- https://www5.cao.go.jp/j-j/wp/wp-je24/csv/d1-1-42.csv
- https://ja.wikipedia.org/wiki/%E3%83%87%E3%82%B8%E3%82%BF%E3%83%AB
- https://www.chisou.go.jp/sousei/pdf/r5_guideline-checkaction.pdf
- https://www.digital.go.jp/
- https://www.pref.yamaguchi.lg.jp/uploaded/attachment/160746.pdf
- https://e-words.jp/w/%E3%83%87%E3%82%B8%E3%82%BF%E3%83%AB.html
- https://www.kantei.go.jp/
- https://dashboard.e-stat.go.jp/
- https://www.kantei.go.jp/jp/kakugikettei/index.html
- https://www.digital.go.jp/resources/japandashboard
- https://www.kantei.go.jp/jp/naikaku/index.html
Share
Related Reports

Code-driven Manifesto: Auditing Local Gov Data and Systems (Kagawa case study)
Local gov systems run but hide data behind UIs; expose CSV/JSON, APIs, and common schemas to unlock value.

Code-driven Check: Japan’s Open Data and the Machine-Readable Gap
Digital Japan has dashboards and rules, but PDFs and messy formats still block automated policy verification; mandate CSV/JSON, APIs, and dataset linting.

Code Speaks: Testing Japan's Gov Data and Dashboards
Japan has great dashboards but inconsistent machine-readability. This report inspects e-Stat, Japan Dashboard, Kantei PDFs, and proposes API-first fixes and practical code examples.