# Sources & method — Vol 1, Issue 2

One section per file in this folder: what the indicator is, which
release it came from, what it does not cover, and anything the figures
should carry as a caveat.

Data freeze: three weeks before publication. After the freeze, revisions
to the underlying series go to the corrections log, not into this file.

## district_access_mpi.csv

**Indicator:** Population-weighted travel time to the nearest health facility
for each of Pakistan's 147 districts, in minutes, on two assumptions, plus the
district's multidimensional poverty index where the poverty survey reaches it.

**Merged districts.** The seven former FATA agencies (Bajaur, Khyber, Kurram, Mohmand, Orakzai, North Waziristan, South Waziristan) are labelled Khyber Pakhtunkhwa in `province` and carry `merged_district = 1`, so every KP figure includes them. The boundary file still draws the pre-2018 agency outlines, which is why they appear as separate polygons on the map. Adaad house rule: refer to them as the ex-FATA merged districts, never as FATA.

- Travel times: Malaria Atlas Project global accessibility layers, motorised
  (2019) and walking-only (2020), 1 km resolution, from Weiss et al., "Global
  maps of travel time to healthcare facilities", *Nature Medicine* 26, 2020 —
  https://www.nature.com/articles/s41591-020-1059-1 and
  https://malariaatlas.org/project-resources/accessibility-to-healthcare/
- Facility set behind those layers: OpenStreetMap and Google Maps hospitals and
  clinics, public and private together. The layers carry no information on
  staffing, opening hours or quality. Travel times are modelled costs to mapped
  destinations, not guaranteed lower bounds on actual journeys.
- Population weights: WorldPop 2020 UN-adjusted 1 km counts, 220.1 million on
  valid pixels — https://hub.worldpop.org/geodata/summary?id=24777
- District boundaries: the 147-district set used across Data Darbar, not
  geoBoundaries ADM2, which is missing Larkana, Chiniot, Nankana Sahib and
  Sujawal and merges Azad Jammu and Kashmir into a single polygon.
- Poverty: the Alkire-Foster multidimensional index Data Darbar constructs from
  PSLM 2019-20 microdata, covering the 119 districts that survey reaches. The
  22 blank cells are survey coverage, not join failures. All seven Karachi
  units inherit the citywide value.

**Caveat carried by the figures:** WorldPop puts Gilgit-Baltistan at about 1.1
million against roughly 1.7 million in the 2023 census, so headcounts for the
sparsely settled north are understated and the region's contribution to the
far-from-care total is larger than these figures show. National shares are the
more robust reading.

**Provenance:** built by `build.py` in `Adaad/data/pk-health-access`, which
clips both rasters, computes the zonal statistics and joins the poverty index.

## health_access_figdata.json

**What it is:** every number the piece's seven figures draw, in one file the
page fetches at runtime. Nothing in the charts is hard-coded into the page, so
a figure cannot drift from the analysis that produced it.

Keys, and where each comes from:

- `national`, `ccdf_mot`, `ccdf_wal`, `districts`, `quintiles` — the travel-time
  distribution, the district scatter and the poverty-fifth gradient, from
  `pk-health-access` as above. Province headcounts beyond an hour keep the
  former tribal areas as their own rows; the text adds them to Khyber
  Pakhtunkhwa (3.85m + 1.53m = 5.4m).
- `gap` — Figure 5: the rural-minus-urban facility-delivery rate across the 90
  mixed districts with complete data, split at the median urban-rural
  walking-time gap, unweighted means (−5.68 and −17.35 points; mean walking
  gaps 51 and 178 minutes). From `outcomes-recalculation.txt` in
  `Adaad/.work/unequal-road-review`, which reproduces `analyze_cells.py`.
- `windows` — Figure 6: the distance coefficient on first-month mortality by
  the births included, district fixed effects, PSU clusters with province
  prefixes. 1–179 months +0.54 (−0.59, +1.66; n 291,600); 1–24 months +2.16
  (−0.10, +4.41; n 39,854); 2–24 months +2.23 (−0.13, +4.59; n 37,986); most
  recent birth only +3.07 (+0.88, +5.26; n 36,429). From
  `births-raw-recalculation.txt` and `delivery-exact-recalculation.txt`,
  recomputed from the raw bh.sav files on 8 September 2026.
- `models` — the other coefficients quoted in the text after the 8 September
  recalculation: delivery −2.07 pp per doubling (−3.18, −0.95; n 36,188, the
  clean home/facility sample); fever +0.55 pp (−0.74, +1.84; n 25,523) with
  wealth −5.87 and +4.90 for the lowest and highest two provincial fifths;
  facility-birth mortality +0.76 per 1,000 (−3.46, +4.99). The earlier
  release quoted about three points for delivery, from a cell-level model,
  and "no detectable gap" for fever from an unweighted demeaning; both are
  corrected here.
- `entry` — Figure 5: share of births in a facility, and the hospital share of
  those, by walking-time band, from `analyze_facility_type.py` in
  `pk-mics-districts`; 37,099 births matched to district urban/rural travel
  cells. Descriptive, unadjusted.
- `risks`, `facility` — the neonatal risk-factor coefficients and the raw
  facility/home mortality rates from the original analysis. No longer drawn
  on the page since the 8 September redraft; kept for the record and for the
  corrections log.

**Sources behind those:** provincial MICS6 microdata with district identifiers,
Punjab 2017-18, Sindh 2018-19, Khyber Pakhtunkhwa 2019, Balochistan 2019-20,
129 districts — https://mics.unicef.org/surveys

**Limits stated in the piece:** the surveys' urban/rural classification and
the map's density threshold (1,500 people per km²) have not been validated
against each other; pooled provincial weights are not nationally
representative; birth histories reach back almost fifteen years against a
2019-20 map; and facility versus home delivery is not an experiment, because
complicated labours select into facilities and the most-recent-birth sample
leaves out earlier deaths. The full specification, including
the standard errors behind every published estimate and the contrasts that are
too noisy to publish, is `SPECIFICATION.md` in `pk-mics-districts`.

**Household fifths:** the MICS wealth index, built from assets, housing
materials, water, sanitation and electricity. MICS records no income, so no
figure in the piece is an income quintile. The population fifths in the
travel-time figures are a different measure, ranking Pakistanis by their
district's multidimensional poverty index.

## Piece 02 · Who pays for the state

**Rebuilt 9 September 2026** on pipeline v2 (revised to v2.2 the same day after
two follow-up reviews; the text is the final editorial redraft of 9 September) after an editorial and methods review found that the shipped income aggregate deducted no farm input
costs, counted self-employed earnings as wages on top of the enterprise
modules, added property-rent detail rows to their total, trusted a business
receipt total that is blank or inconsistent in most households, and applied
the salary schedule to pay the survey records net of tax. Every number in
the piece and both data files changed; the review, the reply to it and the
audit trail are in the analysis folder (`review/who-pays-for-the-state/`).

**Microdata:** HIES 2024–25 (PBS), 30,123 households, field year July
2024–June 2025; LFS 2024–25 for the wage comparison (monthly-paid
employees who reported last month's net cash pay from the main job,
`s7_lfs_compare.py`; LFS bonuses recorded separately and not added). Both
surveys sit on the Census 2023 frame.

**Tax system:** Finance Act 2024 rates as in force during the field year.
GST 18% standard with Sixth Schedule exemptions; medicines at the 1%
concession for registered drugs; stationery and vermicelli at 10% (FBR
Explanatory Circular 03/2024); two-tier cigarette FED; 20% FED on aerated
drinks and packed juices; petroleum levy Rs 60/L to 15 March 2025, Rs 70/L
to 15 April and Rs 78.02/L (petrol) / 77.01 (HSD) after (OGRA), Rs 64.6/L
time-weighted; provincial services taxes on telecom plus 15% withholding;
FY25 salaried PIT schedule with the 10% surcharge above Rs 10m, applied to
pay grossed up from the net amounts HIES records (Section 1B, Note 1);
FY25 non-salaried schedule with the 10% surcharge above Rs 10m for the
statutory scenarios (FBR Circular 01 of 2024-25), person as tax unit. FBR Revenue Division Yearbook 2024–25, Table 9: salary withholding
Rs 605.6bn; payments with returns Rs 221.5bn and collection on demand
Rs 266.7bn across all taxpayer classes, of which Rs 150bn is the piece's
stated assumption for the individual non-salaried share.

**Electricity:** FY25 domestic tariff as notified (NEPRA SRO 1032 of 12
July 2024 and the 11 July 2024 consumer-end tariff decision): lifeline,
protected and unprotected slabs, previous-slab benefit for protected
consumers only, unprotected consumers billed at the slab rate on all units,
fixed charges above 300 units, July–September 2024 relief rates for the
first 200 units. Support valued against NEPRA's FY25 national average
tariff, Rs 35.50/kWh, a regulated average-revenue benchmark, unscaled; the Rs 30 benchmark sensitivity is in the piece's
method note. Protected-consumer definition: non-time-of-use residential connections
(sanctioned load below 5 kW) using 200 units or fewer in each of the
previous six months (NEPRA SRO 1032 of 12 July 2024; K-Electric FAQ). IMF Country Report, third review
(April 2026), RM12: targeted support through BISP by end-January 2027;
Box 2: FY25 tax revenue 12.3% of GDP including provincial taxes and PDL.

**Transfers and comparators:** BISP receipt and amounts from HIES section
8A (code 820); PASS Yearbook 2024–25 for Kafalat disbursement (Rs 469bn).
World Bank ASPIRE indicator `per_sa_allsa.ben_q1_tot`: Pakistan 31.6%
(2018), Bangladesh 25.9% (2022), Brazil 30.2% (2022), retrieved 8
September 2026; all safety nets, welfare quintiles. ONS, *Effects of taxes
and benefits on household income*, FYE 2024, Figure 1 data: cash benefits
including the state pension are 56% of the poorest fifth's disposable
income. Provincial agricultural income tax: IMF first review (May 2025)
for the January 2025 commencement; Sindh Amendment Act XXV of 2025 for the
restoration of the older rates for January–June 2025.

**Files shipped:** `tax_incidence_deciles.csv` (component rates by decile
under both rankings, with the above-benchmark electricity payment, mean
units and household size), `tax_incidence_figdata.json` (all five figures;
`parade` rows carry the same households' median income and median
consumption as fourth and fifth values, `n` carries the audit counts), and
`pipeline/pk-fiscal-incidence/` (the seven-stage pipeline, runnable from
that folder with `HIES_DIR`, `LFS_FILE` and `FI_OUT` set, its README and
changelog, the item-level tax map with every informality judgment, the
block assignment, and the HIES-v-LFS wage percentiles). The same code
lives in the analysis folder as `data/pk-fiscal-incidence/v2`; v1 stays
beside it for provenance.

**Freeze:** analysis re-frozen with the issue on 9 September 2026 after
the rebuild and the final editorial redraft. Later revisions to PBS microdata or FBR collections go to
the corrections log.

## Piece 06 · Social mobility, to a degree (`degree_family_quantiles.csv`, `degree_family_figdata.json`)

Monthly household consumption per adult equivalent at the 5th to 95th
percentile (5-point steps), household-weighted, from the HIES 2024-25
microdata, for households where the head and any spouse recorded as a
member never attended school, with at least one of the head's own sons or
daughters aged 25-40 recorded as a member and no cash remittance reported
(section 8A code 802, or 805-808). Each household is banded by its
best-educated such child: never attended (the questionnaire's own answer),
then highest completed qualification, with pre-primary in primary or
below, diplomas with intermediate, and codes 13-24 as a degree. Households
with any eligible child still enrolled, or with a qualification recorded
as "other", are dropped whole. The `series` column carries "Sons and
daughters" (the figure's default, 1,883 households) and "Sons only"
(behind the toggle, 1,723); the `parents` column carries "Never attended"
(the bars) and "Attended" (the same rules for households whose head or
spouse did attend school; the figure draws only its degree-band median as
the dashed benchmark, 20,282 for the default series across 770 households
and 20,913 for sons only across 652). Adult equivalents (1.0, 0.8 under
18) and the consumption aggregate follow `pk-fiscal-incidence` exactly.

**Caveats that ship with the piece:** the roster is members only, so the
bars describe households whose graduate stayed (temporarily absent members
are included; a child who set up a household elsewhere, or married into
one, has no row); only 15 per cent of resident graduate sons and 18 per
cent of resident graduate daughters have unschooled parents; 87 per cent
of the resident graduate daughters have never married, against 23 per cent
of all graduate women aged 25-40; the degree band is 169 households (131
sons only) with an effective sample of about 134, so its outer percentiles
are soft; the comparison is descriptive, not a return to schooling.
Replaces the two-panel "two ladders" figure after the 9 September source
audit found that section 8A does record the main sender's relation to the
head (code 810), that the matric turning point was an artefact of pooled
bins, and that never-attended had been inferred from zero constructed
years. Built by `Adaad/data/pk-degree-family/build_degree_family.py`.
## Piece 04 · How far is the girls' school? (`school_access_*.csv`, `schools_pk_2026-09.csv.gz`)

**What the files are.** `school_access_districts.csv` is the population-weighted
median distance (km, straight line on a Lambert conformal conic) and walking
time (minutes, Malaria Atlas friction surface) from every populated 1 km cell
to the nearest government school of each sex, for the 132 districts the layer
covers; `gap_km` is girls minus boys. `school_access_district_school_counts.csv`
counts girls' and boys' middle-or-above schools whose point falls inside each
district polygon (Figure 1). `school_access_coverage.csv` is the coverage
ledger: what each region lists, what the layer positions, and where the
positions come from. `school_access_validation_summary.csv` is the scorecard of
the layer against the Mouza Census 2020 (Spearman rank agreement between the
layer's district distances and the villages' own reported distances, by region
and level). `schools_pk_2026-09.csv.gz` is the layer itself: one row per listed
government school (125,317; 121,020 positioned; 118,673 in the analysis) with
`coord_method`, `coord_precision` and `coord_tier` saying how each position was
obtained, `in_analysis` and `analysis_note` saying whether and why a row entered
the analysis, and for Sindh the official SEMIS gender field and enrolment by
sex beside the name designation.

**Canonical copy.** Data Darbar carries the same tables as Parquet with a full
column dictionary and a browser query console
(https://darbar.adaad.org/datasets/schools-pk/); it will carry later releases.
The files here are frozen on the issue date.

**Sources.** SED Balochistan open-data portal (sed.gob.pk, GPS at source);
Sindh SELD Institution Checker roster and Distance Checker
(checker.sindheducation.gov.pk) with RSU district GIS pins, positions
multilaterated on the sphere from 1.2 million checker distances to 540 seed
schools; KP Education Monitoring Authority school locator (roster and
per-school detail) with the JSiMS mirror of KP EMIS for primary schools;
Punjab School Information System roster (sis.pesrp.edu.pk), geocoded from
settlement names against GeoNames and OpenStreetMap; GB EMIS roster, geocoded;
Mirpur and Kotli boards' school rosters, geocoded; OpenStreetMap education
features for Islamabad's federal schools. Population: WorldPop 2020. Enrolment:
PBS Census 2023 district Table 13(b), successor districts summed into 2017
parents. Village reports: PBS Mouza Census 2020 microdata, Form-11 Part IV.

**Coverage.** Government schools only. Settled KP only: the ex-FATA merged
districts (6,394 schools per KP's annual school census) have no public
locations. Two of AJK's ten districts. Islamabad's federal network only (78
schools with a sex designation). The regional grade in `coord_tier` is a grade
of positional quality, not completeness: Punjab is 95 per cent geocoded but only
half of that at settlement precision.

**Definitions.** Networks are by name designation (GG/GB, Girls/Boys). In
Sindh, 23,013 of the schools SEMIS labels are Mixed, and most designated boys'
middle-plus schools enrol girls; the Sindh gap of 0.8 km under designation is
0.1 km if Mixed schools join the girls' network and zero if any school
enrolling girls counts (`girls_km_mixed`, `girls_km_attend` on Data Darbar's
`school_access_district`). `middle_plus` is Middle, Elementary, Secondary or
Higher Secondary (Sindh); Middle, High, Higher Secondary (KP, Punjab, GB,
AJK); M, H, S (Balochistan).

**Validation.** Run on the released layer's `in_analysis` rows. The layer clears
the Mouza Census count floors at middle and high level in every province;
Punjab's primary ratios (0.53 boys', 0.61 girls') reflect private provision, GB's
girls' primary ratio (0.61) reflects sex inferred from names. The ordering of
districts by how far girls are from a middle school agrees with the villages'
(Spearman 0.65 nationally, 0.76 in KP); the sign of the girls-minus-boys gap
agrees in 92–100 per cent of districts in Balochistan, KP and Sindh and in 40
per cent in Punjab, where both sources put the gap within a kilometre of zero.

**Why the positions of girls' schools are published.** Every position is one a
provincial department has itself put online, or follows from distances its own
public distance checker returns, or is a settlement geocoded from a public
roster. The compilation adds nothing beyond the registers: no head teacher,
no telephone number. No attempt was made to locate schools in the merged
districts or the eight AJK districts whose departments publish no list.

## travel_time_tehsils.csv

**What it is:** travel time to the nearest health facility for each of the
553 tehsils in Data Darbar's boundary layer: unweighted and
population-weighted means (minutes, motorized and walking-only), shares of
population beyond 30/60/120 minutes on each surface, WorldPop 2020
population and the count of valid 1 km cells behind every figure. The
Method Notebook's ship file; every number in the piece traces to it or to
the district-level build beside it.

**Merged districts.** The seven former FATA agencies (Bajaur, Khyber, Kurram, Mohmand, Orakzai, North Waziristan, South Waziristan) are labelled Khyber Pakhtunkhwa in `province` and carry `merged_district = 1`, so every KP figure includes them. The boundary file still draws the pre-2018 agency outlines, which is why they appear as separate polygons on the map. Adaad house rule: refer to them as the ex-FATA merged districts, never as FATA.

**Sources:** MAP accessibility surfaces (Weiss et al. 2020, Nature Medicine
26), clipped to Pakistan; WorldPop 2020 UN-adjusted 1 km counts resampled to
the same grid; Data Darbar 553-tehsil boundary layer (2015 vintage), names
joined from the Darbar warehouse.

**Coverage and caveats:** 553 rows, of which 552 have valid cells and positive
population. Manora Cantonment has no valid cells; its blank statistics are
missing estimates, not zero travel time. The tehsil-covered grid represents
1,175,119 cells and 219,576,102 modelled people. Person counts inherit the 2020
base and reporting geography, including AJK and GB; a later national census
total does not update the spatial distribution. Cell count measures spatial
representation, not an independent sample size or a guarantee of reliability.
The population model can shift weighted times in either direction. Facility
coverage and movement assumptions remain uncertainties, not one-sided bounds.

**Provenance:** built by `build_tehsils.py` in `Adaad/data/pk-health-access`,
alongside the district build, using cell-centre boundary assignment and
nearest-neighbour population resampling. The district build covers 1,179,790
cells and 220,111,132 people, so it has 4,671 more cells and 535,030 more people.
Different rasterised boundary coverage changes the denominator; differences
are not simply rounding. The tehsil table recombines to 22.1034 motorized and
127.6976 walking minutes, weighted by population; 203.6183 and 680.1177 weighted
by valid cell counts. More than 60 motorized minutes represents about 16.359M
people (7.4503%) versus about 16.420M in the district output; both round to 16.4M.
Six Killa Saifullah tehsils combine to 112.3 minutes; the district build gives
113.8. Published tehsil means and threshold percentages are rounded to one
decimal, so recombinations are approximate. Thresholds are strictly greater
than 30, 60 or 120 minutes.

**Teaching companions:** `notebook/reproduce_summary.py` recombines the frozen
CSV using Python 3's standard library; `notebook/README.md` contains the exercise,
answers, dictionary and scope of a full GIS rebuild. Both are downloaded under
the issue's data path. They do not rebuild source rasters or new routes.

**Freeze:** analysis frozen with the issue on 9 September 2026.

## Piece 08 · District profiles (`mpi_profile_districts.csv`)

One row each for Khuzdar, Rawalpindi, their provinces and Pakistan. District
rows carry Data Darbar's Alkire-Foster MPI built from PSLM 2019-20 microdata
(index, headcount H, intensity A, sample size, and the seven censored
deprivation headcounts), the PSLM 2019-20 modules (piped water to the
dwelling, no toilet, share who never attended school), and Census 2023
(population, female literacy 10+).

**The pick:** the two ends of the poverty headcount among the 119 districts
whose PSLM sample the build does not flag as small, taking districts rather
than the capital territory. Khuzdar's headcount is the highest outright
(93.8 per cent); Rawalpindi's (9.6) is the lowest after Islamabad's 6.0.
On the composite index Dera Bugti ranks above Khuzdar (0.671 v 0.635)
because its intensity is higher (72.7 v 67.7) — deeper poverty, less of it.
The page states both neighbouring facts rather than hiding them. The poverty
cutoff is the standard Alkire-Foster one third of the weighted indicators,
which is why no district's intensity falls below 41.8.

**Comparators:** province and national figures are population-weighted means
of the district values, weighted by Census 2023 district population, over the
districts where both the MPI build and a census population exist (123 of 141;
PSLM-module comparators cover the districts with that module). The method
reproduces Issue 1's published Punjab headcount (31.7) and the census literacy
aggregates (Punjab 59.8, national 51.9) exactly. The national headcount comes
out at 41.2 against the 41.5 Issue 1 printed; the difference is coverage of
the weight join, and the Issue 1 figure has not been reproduced from the
current warehouse.

**Caveats that ship with the piece:** the MPI is built from the same PSLM
2019-20 microdata as the PSLM rows, so those witnesses are not independent of
each other; Census 2023 is. Khuzdar's two low components are questionnaire
artefacts, and the page carries this as a block, not a footnote: the water
indicator asks whether the source is improved, so wells and hand pumps pass
where nothing is piped (7.0 per cent of Khuzdar households have piped water,
yet water deprivation reads 17.2); the electricity indicator asks whether the
household has a connection, not hours of supply (0.6 per cent). Both biases
understate deprivation where infrastructure exists on paper. No `hies_` field
from the warehouse is used anywhere in the piece: the HIES 2024-25 district
crosswalk misassigns roughly 97 per cent of households and carries no
district identifier for urban households at all.

**Vintage:** the MPI and PSLM figures describe 2019-20; the census figures
describe 2023. Ranks quoted on the page are among the 119 full-sample
districts, rank 1 most deprived.

**Freeze:** analysis frozen with the issue on 9 September 2026.

## middle_budget_categories.csv and middle_budget_figdata.json

**Indicator:** Budget shares of twelve expenditure categories for households in
the national third consumption quintile, HIES 2018-19 and HIES 2024-25, with
the change in each category's real spending per person between the rounds.

- Source: PBS Household Integrated Economic Survey microdata, 2018-19 and
  2024-25 — https://www.pbs.gov.pk/ — harmonised item by item into common
  categories by Ghufran Khalid; 5,008 and 5,884 households in the quintile.
- Quintile: households ranked by per-capita consumption, population weights,
  separately in each round.
- Deflation: PBS group-wise CPI matched by interview month and urban or rural
  residence, 2015-16 prices — https://www.pbs.gov.pk/price-statistics
- Necessities are food and non-alcoholic beverages; housing, water,
  electricity, gas and fuels; health.

**Values.** All figures are the author's cleaned aggregates delivered on
6 September 2026 (Aggregate Data Master File and Chart Data). Real spending
per person at 2015-16 prices: total 4,105 to 3,887 (−5.3 per cent; −9.3 per
household, mean size 6.41 to 6.14); food 1,782 to 1,563 (−12.3); housing and
utilities 819 to 975 (+19.0); health 130 to 137 (+4.8); the three necessities
2,732 to 2,674 (−2.1); everything else 1,373 to 1,213 (−11.7). Food plus
health −11.2; everything except housing −11.4. Communication is in the
shares table but has no plotted real change: +60.1 per cent under the broad
item match, +109.8 under the strict one.

**Alternative housing deflator.** The author's component-level deflation
puts housing and utilities at 828 to 1,051 (+26.8). Substituting those
levels and holding other categories fixed gives necessities +0.3 and total
−3.7. Open with the author: how rent (actual and imputed) was reported in
each round and how the components were matched to the CPI sub-indices.

**Provinces.** Point estimates for the necessities share rose in all four;
the supplied intervals include zero for Balochistan (share and real change)
and Khyber Pakhtunkhwa (share). Confidence level not stated in the file;
notes describe an approximate PSU bootstrap.

**Boundary check.** 4,645 of 5,008 and 5,450 of 5,884 households retained
(about 93 per cent of the sample, unweighted); the weighted retention is
still to come from the author.
