Code Speaks: Auditing Japan’s Public CSVs and Data Practice

IT Policy Proposals
Code Speaks: Auditing Japan’s Public CSVs and Data Practice

どうも〜おかむーです! Today I'm diving into some public CSVs from Japanese government sites with an engineer's eye — "code talks" style. This is a technical take on machine-readability, APIs, and how policy claims map to data.

  • This article inspects a few live CSVs (new-car sales, public investment, hospital list, NICTER notices).
  • Main finding: data is published, but machine-friendliness, metadata, and APIs are inconsistent — lots of low-hanging UX fixes.
  • I give concrete parsing tips, a short Python recipe, and actionable improvements (OpenAPI, CSVW/JSON-LD, catalogs).

Conclusion

Public bodies are publishing useful CSVs (see https://www5.cao.go.jp/j-j/wp/wp-je24/csv/d1-1-4.csv and https://www5.cao.go.jp/j-j/wp/wp-je24/csv/d1-1-42.csv, plus https://www.digital.go.jp/assets/.../xxxxxx_hospital.csv and https://notice.go.jp/docs/status_nicter.csv). But from a developer standpoint, the work stops at dumping files: missing stable APIs, ambiguous encodings/units, absent schema/metadata, and no discoverable machine contracts. 要するに、データはあるけど“API-first”の体験には程遠いということです。

Report

Data sources I eyeballed

  • New car sales timeseries (cabinet office): https://www5.cao.go.jp/j-j/wp/wp-je24/csv/d1-1-4.csv
  • Public investment budget vs settlement: https://www5.cao.go.jp/j-j/wp/wp-je24/csv/d1-1-42.csv
  • Hospital registry CSV (Digital Agency asset): https://www.digital.go.jp/.../xxxxxx_hospital.csv
  • NICTER notice feed: https://notice.go.jp/docs/status_nicter.csv

Look at this: some CSVs contain a BOM and Japanese headers — fine for humans, but code needs to know encoding and units!

Technical assessment (quick)

  • Formats: Mostly CSV — great for simplicity, but CSV alone lacks typed schema and units. PDF-to-CSV conversions sometimes leave artifacts.
  • Encoding: BOM/Shift_JIS/UTF-8 ambiguity observed in snippets. エンジニア的に言うと, you will hit decode errors unless you handle encodings explicitly.
  • Metadata: No machine-readable schema (CSVW/JSON-LD) attached. No clear unit fields (e.g., is sales in thousands or ten-thousands?).
  • API: No documented REST/GraphQL API for these datasets. That makes live apps fragile (screen-scrape + CSV polling).
  • Provenance & licensing: Often implicit; needs explicit license & update cadence fields.

Quick reproducible parsing recipe (Python)

import requests

import io

import pandas as pd

url = 'https://www5.cao.go.jp/j-j/wp/wp-je24/csv/d1-1-4.csv'

res = requests.get(url)

Try utf-8-sig to handle BOM; fall back to shift_jis if needed

data = io.StringIO(res.content.decode('utf-8-sig'))

df = pd.read_csv(data)

print(df.head())

parsing tips: rename columns to ascii keys, add units column based on metadata lookup

Notes: handle utf-8-sig, try shift_jis if decode fails; coerce numeric with pd.to_numeric(errors='coerce').

Policy-to-data mapping: how to check targets vs reality

  • Ensure you know units & aggregation level (monthly, 3MA, totals).
  • Parse budget vs settlement fields, convert to common base (JPY millions).
  • Compute metrics: gap = (settlement - budget) / budget. Visualize trend and CI.
  • This workflow lets you verify claims like “public investment increased X%” against machine data.

    Concrete improvement proposals

    • Publish OpenAPI-backed JSON endpoints for each dataset, with stable versioning.
    • Attach CSVW metadata (tableSchema) or JSON Schema + example payloads. Include: units, frequency, last_updated, contact, license.
    • Standardize encoding to UTF-8 (no BOM) and UTF-8 Content-Type headers.
    • Provide dataset catalog (CKAN or DataPortal) with machine-searchable tags; link to Japan Dashboard/e-Stat.
    • Ship example ETL scripts (Python/R) in a public GitHub repo so civictech teams can contribute.
    • Offer webhooks or push-notifications for update events (or simple ETag/Last-Modified headers).

    まとめ

    • Good: government publishes raw CSVs — that's the baseline win!
    • Bad: missing metadata, inconsistent encodings, and no API contract make rebuilding dashboards fragile.
    • Actionable: add schema, standardize UTF-8, publish OpenAPI/CSVW, and provide example code — these are high-ROI fixes.

    おかむーから一言

    As an entrepreneur-engineer, I want gov data to be as easy to import as pip install. Let's make data discoverable, verifiable, and API-first — code should be the manifest of policy, not a guessing game.