Python Analytics
Nafs Results Analytics — Aseer Region
A complete Python pipeline that starts from the raw extract, applies data-quality rules and quarantines failures, then builds an analytical model of trends, gaps and school segments that feeds this executive dashboard.
- 1Raw extract70,041 rows
- 2DQ rulesDQ-001…007
- 3Modelpandas · SciPy
- 4DashboardJSON → Chart.js
Key findings
Generated by the pipeline- Regional proficiency rose from 32.0% (1444 AH) to 40.4% (1447 AH), +8.4 pts at 2.9 pts per year.
- Gap to the national average is -1.8 pts in 1447 AH vs -3.0 in 1444 AH.
- Science leads (52.2%) and Mathematics trails (31.5%), a 20.7-pt spread; weakest sub-domain: Algebra (26.9%).
- Top governorate Abha (45.3%), lowest Al-Harjah (29.5%); most improved Tanomah (+5.0).
- Girls' schools outperform by 5.5 pts on average; the difference is statistically significant (Welch t-test, p < 0.001).
- 189 schools flagged for priority intervention: ≥5 pts below the regional average and down ≥5 pts year-on-year (of 2,137 classified).
- On the current trend, 1448 AH proficiency is projected at ~43.4% (80% range: 42.5–44.2%).
Proficiency by subject & grade
Performance-level distribution
Share of students per levelGovernorate ranking
Dashed line: regional averageGender & authority
Sub-domain diagnosis
Grade 6 science & Grade 9 mathsSchool segmentation: level vs change
Each dot is a school · 1447 AH vs 1446 · min. 15 testedPriority schools for intervention
≥5 pts below average and down ≥5 pts · pseudonymised IDs| School | Governorate | 1447 | Δ |
|---|---|---|---|
SCH-1727 · Girls |
Ahad Rufaidah | 8.9% | -22.4 |
SCH-0770 · Boys |
Bisha | 16.7% | -23.9 |
SCH-1801 · Boys |
Bariq | 11.1% | -16.2 |
SCH-1822 · Girls |
Al-Majardah | 20.7% | -24.1 |
SCH-1592 · Boys |
Balqarn | 18.5% | -21.5 |
SCH-1631 · Boys |
Balqarn | 10.3% | -12.7 |
SCH-1337 · Girls |
Tathlith | 5.9% | -7.9 |
SCH-2114 · Boys |
Al-Harjah | 8.0% | -9.9 |
SCH-0016 · Boys |
Khamis Mushait | 9.1% | -10.9 |
SCH-1020 · Girls |
Muhayil | 6.9% | -7.4 |
SCH-1684 · Girls |
Ahad Rufaidah | 14.8% | -15.2 |
SCH-1625 · Boys |
Balqarn | 14.8% | -14.8 |
Total flagged: 189 · Full list in priority_schools.csv
Data quality before analysis
Raw-extract quality scorecard
The extract is checked against seven rules mapped to DAMA dimensions before any calculation. Failures are quarantined or conformed from reference data, with the action documented.
Completeness
99.99%
Consistency
100.00%
Validity
100.00%
Uniqueness
99.99%
Timeliness
100.00%
Scale 95–100% · dark tick = target
| Rule | Dimension | Check | Checked | Failed | Pass | Action |
|---|---|---|---|---|---|---|
DQ-001 | Completeness | School ID is not null | 70,041 | 5 | 99.99% | Quarantine & return to source |
DQ-002 | Consistency | Governorate matches reference data | 70,041 | 2 | 100.00% | Auto-conform from reference table |
DQ-003 | Validity | Grade & domain codes are in reference lists | 70,041 | 1 | 100.00% | Quarantine & fix coding |
DQ-004 | Validity | Proficiency within 0-100 and consistent with counts | 70,041 | 6 | 99.99% | Recompute from source counts |
DQ-005 | Validity | Tested does not exceed expected | 70,041 | 3 | 100.00% | Quarantine for source verification |
DQ-006 | Uniqueness | No duplicate year/school/grade/domain | 70,041 | 9 | 99.99% | Remove duplicates after review |
DQ-007 | Timeliness | Extract covers all reporting years | 4 | 0 | 100.00% | Escalate delay to data owner |
Method & statistics
Behind the numbers
Girls vs boys schools+5.5 pts
Welch t = 8.58 · p < 0.001 · n = 1,112 / 1,025
Private & intl. vs public+4.6 pts
Welch t = 4.07 · p < 0.001
School size vs proficiencyρ = -0.005
Spearman · p = 0.817 — no material relationship
Annual improvement+2.9 pts/yr
OLS · R² = 0.996
Projection for 1448 AH43.4%
80% prediction interval: 42.5–44.2%
Pipeline excerpt
nafs_pipeline.py# DQ-004 validity — proficiency within 0–100 and consistent with counts
recomputed = 100 * df["proficient"] / df["tested"]
bad_rate = (df["prof_rate"] < 0) | (df["prof_rate"] > 100) \
| ((recomputed - df["prof_rate"]).abs() > 0.5)
rule("DQ-004", "Validity", bad_rate.sum(), action="Recompute from source counts")
# Priority = materially below average AND materially declining
wide["priority"] = ((wide["rate_1447"] <= region_avg - MATERIAL)
& (wide["change"] <= -MATERIAL)
& (wide["tested_1447"] >= PRIORITY_N))
# Significance: girls' vs boys' schools (Welch t-test)
t = stats.ttest_ind(girls, boys, equal_var=False)
- DataSchool-level extract in the Nafs report layout: year, grade, domain & sub-domain, tested, proficient, performance levels and national average (1444–1447 AH).
- PrivacySchool identifiers are pseudonymised; school-level indicators are suppressed below 15 tested students.
- MeasuresProficiency = proficient ÷ tested (weighted by tested), because the score scale changed across years.
Static Python outputs for reports
Data note: figures on this page are generated from a modelled dataset that follows the Nafs report layout; the same pipeline runs unchanged on the official extract.