← All posts

What building StatLab Zim taught me about “properly”

2026-09-03Career reflection4 min read

I wanted to build StatLab Zim properly. Not perfectly — properly. There's a difference.

It's a Streamlit app for introductory statistics. Upload a CSV, get descriptives, visualizations, and hypothesis tests. Sounds simple. The hard part wasn't the statistics — it was the empty file with only a header.

The first properly moment

On day one, pandas.read_csv crashed on an empty file. No rows, no columns. My data-quality summary tried to count inferred types on an empty DataFrame and threw a KeyError. The user saw a red traceback instead of a helpful message.

Properly meant: an empty CSV should show "Header but no data rows. Visualizations will be limited." — not a stack trace. I added a guard:

if type_summary.empty:
    n_numeric = n_categorical = 0
    # …

Small change, big kindness.

Properly is not about adding features. It's about handling the boring edge so the human on the other side feels respected.

19 tests later

I wrote 19 pytest tests before I felt comfortable shipping: row/col counts, missing-value %, duplicate detection, numeric detection, descriptive spot-checks with pytest.approx, and — crucially — empty/invalid file handling. The tests caught a string-dtype bug where is_object_dtype returned False for pandas 2's new StringDtype.

In Harare we say "Chitsva charimutsoka" — what's new is in the feet; you learn by walking. The tests were the walking.

Shipping beats polishing

I could have added a database, user accounts, a dark mode. I chose light-only for the blog and a restrained palette for the app. Light is honest; you can't hide bad spacing.

StatLab Zim is live at statlab-zim.streamlit.app and the code is at github.com/alfredshingai/statlab-zim. It still has rough edges — kaleido needs Chrome for PNG exports. But it handles an empty CSV gracefully. That's properly enough for today.

— Alfred, second coffee, probably.