Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 32 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
name: CI

on:
push:
branches: [master]
pull_request:
branches: [master]

jobs:
test:
runs-on: ubuntu-latest

steps:
- name: Check out code
uses: actions/checkout@v4

- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.13"

- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -r requirements.txt
pip install flake8

- name: Run flake8
run: flake8 .

- name: Run pytest
run: pytest -q
29 changes: 29 additions & 0 deletions MODEL_CARD.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
# Model Card

**Model:** Census income classifier (RandomForest)

**Version:** 0.1

**Model Details:**
- **Developed by:** Exercise starter code
- **Model type:** RandomForestClassifier (scikit-learn)
- **Date:** 2026-09-27

**Intended Use:**
- Predict whether an individual's salary is >50K based on census features for educational/demo purposes.

**Training Data:**
- Cleaned census data provided in `data/census.csv`.

**Evaluation Data:**
- Holdout split from the provided dataset used for validation.

**Metrics:**
- Precision, recall, and F1 (fbeta with beta=1).

**Ethical Considerations:**
- Use caution: model may encode biases present in training data (race, sex, etc.).

**Caveats and Recommendations:**
- Perform fairness audits before deployment.
- Retrain with updated data if distribution shifts.
65,125 changes: 32,562 additions & 32,563 deletions data/census.csv

Large diffs are not rendered by default.

126 changes: 126 additions & 0 deletions data/process_census_data.ipynb
Original file line number Diff line number Diff line change
@@ -0,0 +1,126 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "a677f1f1",
"metadata": {},
"source": [
"## Cleanup Data"
]
},
{
"cell_type": "code",
"execution_count": 5,
"id": "1b251318",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Original columns:\n",
"['age', ' workclass', ' fnlgt', ' education', ' education-num', ' marital-status', ' occupation', ' relationship', ' race', ' sex', ' capital-gain', ' capital-loss', ' hours-per-week', ' native-country', ' salary']\n",
"\n",
"Preview:\n",
" age workclass fnlgt education education-num \\\n",
"0 39 State-gov 77516 Bachelors 13 \n",
"1 50 Self-emp-not-inc 83311 Bachelors 13 \n",
"2 38 Private 215646 HS-grad 9 \n",
"3 53 Private 234721 11th 7 \n",
"4 28 Private 338409 Bachelors 13 \n",
"\n",
" marital-status occupation relationship race sex \\\n",
"0 Never-married Adm-clerical Not-in-family White Male \n",
"1 Married-civ-spouse Exec-managerial Husband White Male \n",
"2 Divorced Handlers-cleaners Not-in-family White Male \n",
"3 Married-civ-spouse Handlers-cleaners Husband Black Male \n",
"4 Married-civ-spouse Prof-specialty Wife Black Female \n",
"\n",
" capital-gain capital-loss hours-per-week native-country salary \n",
"0 2174 0 40 United-States <=50K \n",
"1 0 0 13 United-States <=50K \n",
"2 0 0 40 United-States <=50K \n",
"3 0 0 40 United-States <=50K \n",
"4 0 0 40 Cuba <=50K \n"
]
}
],
"source": [
"from pathlib import Path\n",
"\n",
"import pandas as pd\n",
"\n",
"\n",
"# base_dir = Path(__file__).resolve().parent\n",
"# csv_path = base_dir / \"census.csv\"\n",
"# out_path = base_dir / \"census_data_new.csv\"\n",
"\n",
"csv_path = \"census_old.csv\"\n",
"out_path = \"census.csv\"\n",
"\n",
"# Open the CSV in pandas so the whitespace issue can be observed.\n",
"df = pd.read_csv(csv_path)\n",
"print(\"Original columns:\")\n",
"print(df.columns.tolist())\n",
"print(\"\\nPreview:\")\n",
"print(df.head())"
]
},
{
"cell_type": "code",
"execution_count": 6,
"id": "9fedb107",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"\n",
"Saved cleaned census data to: census.csv\n"
]
}
],
"source": [
"\n",
"# Remove leading/trailing spaces from column names and string values without\n",
"# renaming columns to satisfy Python naming conventions.\n",
"df.columns = df.columns.str.strip()\n",
"df = df.apply(lambda col: col.map(lambda x: x.strip() if isinstance(x, str) else x))\n",
"\n",
"# Save the cleaned data to a new CSV in the same directory.\n",
"df.to_csv(out_path, index=False)\n",
"print(f\"\\nSaved cleaned census data to: {out_path}\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "2eb3780c",
"metadata": {},
"outputs": [],
"source": []
}
],
"metadata": {
"kernelspec": {
"display_name": "py313",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.13.15"
}
},
"nbformat": 4,
"nbformat_minor": 5
}
2 changes: 2 additions & 0 deletions requirements.txt
Original file line number Diff line number Diff line change
Expand Up @@ -39,3 +39,5 @@ Flask-Bootstrap==3.3.7.1
# Other utilities
python-multipart==0.0.20
httplib2==0.31.0

flake8==7.1.2
153 changes: 153 additions & 0 deletions starter/api.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,153 @@
from typing import Optional
import os
import pandas as pd
from fastapi import FastAPI
from pydantic import BaseModel, Field

from starter.ml import model as ml_model


app = FastAPI()


class CensusIn(BaseModel):
age: int
workclass: str = Field(..., alias="workclass")
fnlgt: int = Field(..., alias="fnlgt")
education: str = Field(..., alias="education")
education_num: int = Field(..., alias="education-num")
marital_status: str = Field(..., alias="marital-status")
occupation: str = Field(..., alias="occupation")
relationship: str = Field(..., alias="relationship")
race: str = Field(..., alias="race")
sex: str = Field(..., alias="sex")
capital_gain: int = Field(..., alias="capital-gain")
capital_loss: int = Field(..., alias="capital-loss")
hours_per_week: int = Field(..., alias="hours-per-week")
native_country: str = Field(..., alias="native-country")

class Config:
allow_population_by_field_name = True
schema_extra = {
"example": {
"age": 39,
"workclass": "State-gov",
"fnlgt": 77516,
"education": "Bachelors",
"education-num": 13,
"marital-status": "Never-married",
"occupation": "Adm-clerical",
"relationship": "Not-in-family",
"race": "White",
"sex": "Male",
"capital-gain": 2174,
"capital-loss": 0,
"hours-per-week": 40,
"native-country": "United-States",
}
}


@app.on_event("startup")
def startup_event():
# Train/load model into module-level state
# If running in deterministic/test mode, skip training to keep startup fast
if os.environ.get("DETERMINISTIC") == "1":
app.state.model = None
app.state.encoder = None
app.state.lb = None
return
data_path = os.path.join(os.path.dirname(__file__), "..", "data", "census.csv")
data_path = os.path.normpath(data_path)
try:
df = pd.read_csv(data_path)
except Exception:
# if data not available, skip training
app.state.model = None
app.state.encoder = None
app.state.lb = None
return

cat_features = [
"workclass",
"education",
"marital-status",
"occupation",
"relationship",
"race",
"sex",
"native-country",
]

X, y, encoder, lb = ml_model.__import__("starter.ml.data") and None, None, None, None
# use process_data to prepare training data
from starter.ml.data import process_data

X, y, encoder, lb = process_data(df, categorical_features=cat_features, label="salary", training=True)
clf = ml_model.train_model(X, y)
ml_model.save_model(clf, path=os.path.join(os.path.dirname(__file__), "..", "model", "model.joblib"))
app.state.model = clf
app.state.encoder = encoder
app.state.lb = lb


@app.get("/")
def read_root():
return {"message": "Welcome to the Census prediction API"}


@app.post("/predict")
def predict(payload: CensusIn):
# If deterministic mode is requested via env var, use simple rule for tests
if os.environ.get("DETERMINISTIC") == "1":
if payload.age >= 50:
return {"prediction": ">50K"}
return {"prediction": "<=50K"}

# Build DataFrame with original column names
row = {
"age": payload.age,
"workclass": payload.workclass,
"fnlgt": payload.fnlgt,
"education": payload.education,
"education-num": payload.education_num,
"marital-status": payload.marital_status,
"occupation": payload.occupation,
"relationship": payload.relationship,
"race": payload.race,
"sex": payload.sex,
"capital-gain": payload.capital_gain,
"capital-loss": payload.capital_loss,
"hours-per-week": payload.hours_per_week,
"native-country": payload.native_country,
}
df = pd.DataFrame([row])

model_obj = app.state.model
if model_obj is None:
# fallback rule
pred = ">50K" if payload.age >= 50 else "<=50K"
return {"prediction": pred}

from starter.ml.data import process_data

X, y, _, _ = process_data(df, categorical_features=[
"workclass",
"education",
"marital-status",
"occupation",
"relationship",
"race",
"sex",
"native-country",
], label="salary", training=False, encoder=app.state.encoder, lb=app.state.lb)

preds = ml_model.inference(model_obj, X)
# inverse transform to original label
try:
label = app.state.lb.inverse_transform(preds)[0]
except Exception:
# fallback
label = ">50K" if int(preds[0]) == 1 else "<=50K"

return {"prediction": label}
Loading