Nafs Results Analytics — Aseer Region¶
Adel Asiri · CDMP · PMP
Pipeline: raw extract → data-quality rules → conformed dataset → analytical model → executive dashboard feed.
| Step | Output |
|---|---|
| 1. Load & profile | raw extract + reference data |
| 2. Data quality (DQ-001…DQ-007) | scorecard, quarantine |
| 3. Model | main-domain & sub-domain tables |
| 4. Analysis | trends, gaps, segmentation, statistics, forecast |
| 5. Export | dashboard.json, charts, CSVs |
In [1]:
import pandas as pd, numpy as np
import nafs_pipeline as P
pd.set_option('display.width', 160)
raw, govs, subjects, grades = P.load()
print(f'{len(raw):,} rows × {raw.shape[1]} columns')
raw.head()
70,041 rows × 23 columns
Out[1]:
| year | العام الدراسي - ميلادي | school_id | إدارة التعليم | المنطقة | authority | gender | gov | stage | حجم المدرسة | ... | expected | tested | proficient | prof_rate | score | lvl_high | lvl_mid | lvl_low | lvl_vlow | national | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 1446 | 2024-2025 | 900614.0 | الإدارة العامة للتعليم بمنطقة عسير | عسير | حكومي | بنات | أبها | ابتدائية | كبيرة | ... | 90 | 81 | 64 | 79.01 | NaN | NaN | NaN | NaN | NaN | NaN |
| 1 | 1444 | 2022-2023 | 900630.0 | الإدارة العامة للتعليم بمنطقة عسير | عسير | حكومي | بنات | أبها | ابتدائية | متوسطة | ... | 29 | 26 | 19 | 73.08 | 67.117 | 73.08 | 18.49 | 4.14 | 4.29 | 31.7 |
| 2 | 1446 | 2024-2025 | 901687.0 | الإدارة العامة للتعليم بمنطقة عسير | عسير | حكومي | بنين | أحد رفيدة | متوسطة | كبيرة | ... | 60 | 58 | 18 | 31.03 | NaN | NaN | NaN | NaN | NaN | NaN |
| 3 | 1447 | 2025-2026 | 900210.0 | الإدارة العامة للتعليم بمنطقة عسير | عسير | حكومي | بنات | خميس مشيط | ابتدائية | كبيرة | ... | 81 | 75 | 25 | 33.33 | 61.285 | 33.33 | 29.09 | 26.57 | 11.01 | 54.9 |
| 4 | 1447 | 2025-2026 | 900418.0 | الإدارة العامة للتعليم بمنطقة عسير | عسير | حكومي | بنات | أبها | ابتدائية | كبيرة | ... | 85 | 77 | 46 | 59.74 | NaN | NaN | NaN | NaN | NaN | NaN |
5 rows × 23 columns
1 · Profiling¶
In [2]:
prof = pd.DataFrame({'dtype': raw.dtypes.astype(str), 'nulls': raw.isna().sum(), 'distinct': raw.nunique()})
prof
Out[2]:
| dtype | nulls | distinct | |
|---|---|---|---|
| year | int64 | 0 | 4 |
| العام الدراسي - ميلادي | str | 0 | 4 |
| school_id | float64 | 5 | 2146 |
| إدارة التعليم | str | 0 | 1 |
| المنطقة | str | 0 | 1 |
| authority | str | 0 | 3 |
| gender | str | 0 | 2 |
| gov | str | 0 | 19 |
| stage | str | 0 | 3 |
| حجم المدرسة | str | 0 | 3 |
| grade | int64 | 0 | 4 |
| domain | str | 0 | 3 |
| subdomain | str | 0 | 10 |
| expected | int64 | 0 | 102 |
| tested | int64 | 0 | 99 |
| proficient | int64 | 0 | 79 |
| prof_rate | float64 | 0 | 2031 |
| score | float64 | 31484 | 23134 |
| lvl_high | float64 | 31484 | 1748 |
| lvl_mid | float64 | 31484 | 5127 |
| lvl_low | float64 | 31484 | 4024 |
| lvl_vlow | float64 | 31484 | 2851 |
| national | float64 | 31484 | 30 |
2 · Data-quality rules¶
Each rule maps to a DAMA dimension, has an owner action, and failed records are quarantined rather than silently dropped.
In [3]:
clean, dq, dims, quarantine, summary = P.data_quality(raw, govs, subjects, grades)
print(summary)
dq[['id','dim_en','rule_en','checked','failed','pass_rate','action_en']]
{'raw_rows': 70041, 'clean_rows': 70023, 'quarantined': 9, 'duplicates_removed': 9, 'conformed': 2, 'overall': 100.0}
Out[3]:
| id | dim_en | rule_en | checked | failed | pass_rate | action_en | |
|---|---|---|---|---|---|---|---|
| 0 | DQ-001 | Completeness | School ID is not null | 70041 | 5 | 99.99 | Quarantine & return to source |
| 1 | DQ-002 | Consistency | Governorate matches reference data | 70041 | 2 | 100.00 | Auto-conform from reference table |
| 2 | DQ-003 | Validity | Grade & domain codes are in reference lists | 70041 | 1 | 100.00 | Quarantine & fix coding |
| 3 | DQ-004 | Validity | Proficiency within 0-100 and consistent with c... | 70041 | 6 | 99.99 | Recompute from source counts |
| 4 | DQ-005 | Validity | Tested does not exceed expected | 70041 | 3 | 100.00 | Quarantine for source verification |
| 5 | DQ-006 | Uniqueness | No duplicate year/school/grade/domain | 70041 | 9 | 99.99 | Remove duplicates after review |
| 6 | DQ-007 | Timeliness | Extract covers all reporting years | 4 | 0 | 100.00 | Escalate delay to data owner |
In [4]:
dims
Out[4]:
| dim_ar | dim_en | score | target | |
|---|---|---|---|---|
| 0 | الاكتمال | Completeness | 99.99 | 99.5 |
| 1 | الاتساق | Consistency | 100.00 | 99.5 |
| 2 | الصحة | Validity | 100.00 | 99.0 |
| 3 | التفرد | Uniqueness | 99.99 | 100.0 |
| 4 | الحداثة | Timeliness | 100.00 | 100.0 |
In [5]:
quarantine[['dq_rule','year','school_id','gov','grade','domain','tested','expected','prof_rate']]
Out[5]:
| dq_rule | year | school_id | gov | grade | domain | tested | expected | prof_rate | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | DQ-001 | 1447 | NaN | سراة عبيدة | 6 | العلوم | 12 | 14 | 33.33 |
| 1 | DQ-001 | 1444 | NaN | أبها | 3 | الرياضيات | 26 | 29 | 53.85 |
| 2 | DQ-001 | 1445 | NaN | خميس مشيط | 6 | العلوم | 31 | 31 | 35.48 |
| 3 | DQ-001 | 1446 | NaN | بيشة | 6 | العلوم | 28 | 33 | 78.57 |
| 4 | DQ-001 | 1444 | NaN | رجال ألمع | 6 | الرياضيات | 12 | 13 | 33.33 |
| 5 | DQ-003 | 1444 | 901323.0 | تثليث | 7 | العلوم | 28 | 31 | 46.43 |
| 6 | DQ-005 | 1446 | 900496.0 | أبها | 9 | الرياضيات | 74 | 67 | 32.43 |
| 7 | DQ-005 | 1447 | 902029.0 | النماص | 6 | العلوم | 15 | 8 | 40.00 |
| 8 | DQ-005 | 1444 | 900589.0 | أبها | 6 | العلوم | 33 | 26 | 30.30 |
3 · Analytical model¶
In [6]:
main, subs = P.build_model(clean, subjects)
res = P.analyse(main, subs, govs)
pd.DataFrame(res['trend'])
Out[6]:
| year | region | national | participation | |
|---|---|---|---|---|
| 0 | 1444 | 31.97 | 34.94 | 94.33 |
| 1 | 1445 | 34.59 | 37.15 | 94.45 |
| 2 | 1446 | 37.98 | 41.04 | 94.40 |
| 3 | 1447 | 40.38 | 42.14 | 94.47 |
4 · Findings¶
In [7]:
for i in res['insights']:
print('•', i['en'])
• Regional proficiency rose from 32.0% (1444 AH) to 40.4% (1447 AH), +8.4 pts at 2.9 pts per year. • Gap to the national average is -1.8 pts in 1447 AH vs -3.0 in 1444 AH. • Science leads (52.2%) and Mathematics trails (31.5%), a 20.7-pt spread; weakest sub-domain: Algebra (26.9%). • Top governorate Abha (45.3%), lowest Al-Harjah (29.5%); most improved Tanomah (+5.0). • Girls' schools outperform by 5.5 pts on average; the difference is statistically significant (Welch t-test, p < 0.001). • 189 schools flagged for priority intervention: ≥5 pts below the regional average and down ≥5 pts year-on-year (of 2,137 classified). • On the current trend, 1448 AH proficiency is projected at ~43.4% (80% range: 42.5–44.2%).
In [8]:
pd.Series(res['statistics']['gender']), pd.Series(res['statistics']['forecast'])
Out[8]:
(girls 4.308000e+01 boys 3.756000e+01 diff 5.520000e+00 t 8.580000e+00 p 1.817803e-17 n_girls 1.112000e+03 n_boys 1.025000e+03 dtype: float64, year 1448.0 value 43.4 low 42.5 high 44.2 dtype: float64)
In [9]:
pd.DataFrame(res['priority']).head(10)
Out[9]:
| id | gov | gov_en | stage | gender | rate | prev | change | tested | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | SCH-1727 | أحد رفيدة | Ahad Rufaidah | ابتدائية | بنات | 8.9 | 31.2 | -22.4 | 45 |
| 1 | SCH-0770 | بيشة | Bisha | مجمع | بنين | 16.7 | 40.5 | -23.9 | 72 |
| 2 | SCH-1801 | بارق | Bariq | متوسطة | بنين | 11.1 | 27.3 | -16.2 | 27 |
| 3 | SCH-1822 | المجاردة | Al-Majardah | ابتدائية | بنات | 20.7 | 44.8 | -24.1 | 58 |
| 4 | SCH-1592 | بلقرن | Balqarn | متوسطة | بنين | 18.5 | 40.0 | -21.5 | 27 |
| 5 | SCH-1631 | بلقرن | Balqarn | ابتدائية | بنين | 10.3 | 23.1 | -12.7 | 58 |
| 6 | SCH-1337 | تثليث | Tathlith | مجمع | بنات | 5.9 | 13.8 | -7.9 | 101 |
| 7 | SCH-2114 | الحرجة | Al-Harjah | متوسطة | بنين | 8.0 | 17.9 | -9.9 | 75 |
| 8 | SCH-0016 | خميس مشيط | Khamis Mushait | متوسطة | بنين | 9.1 | 20.0 | -10.9 | 33 |
| 9 | SCH-1020 | محايل | Muhayil | ابتدائية | بنات | 6.9 | 14.3 | -7.4 | 58 |
5 · Charts & export¶
In [10]:
P.charts(res, dims, govs, subjects)
feed = P.export(res, dq, dims, summary, govs, subjects, grades)
from IPython.display import Image, display
for f in sorted((P.CHARTS).glob('*.png')):
display(Image(filename=str(f), width=640))