What building StatLab Zim taught me about “properly”
I wanted to build StatLab Zim properly. Not perfectly — properly. There's a difference.
It's a Streamlit app for introductory statistics. Upload a CSV, get descriptives, visualizations, and hypothesis tests. Sounds simple. The hard part wasn't the statistics — it was the empty file with only a header.
The first properly moment
On day one, pandas.read_csv crashed on an empty file. No rows, no columns. My data-quality summary tried to count inferred types on an empty DataFrame and threw a KeyError. The user saw a red traceback instead of a helpful message.
Properly meant: an empty CSV should show "Header but no data rows. Visualizations will be limited." — not a stack trace. I added a guard:
if type_summary.empty:
n_numeric = n_categorical = 0
# …
Small change, big kindness.
Properly is not about adding features. It's about handling the boring edge so the human on the other side feels respected.
19 tests later
I wrote 19 pytest tests before I felt comfortable shipping: row/col counts, missing-value %, duplicate detection, numeric detection, descriptive spot-checks with pytest.approx, and — crucially — empty/invalid file handling. The tests caught a string-dtype bug where is_object_dtype returned False for pandas 2's new StringDtype.
In Harare we say "Chitsva charimutsoka" — what's new is in the feet; you learn by walking. The tests were the walking.
Shipping beats polishing
I could have added a database, user accounts, a dark mode. I chose light-only for the blog and a restrained palette for the app. Light is honest; you can't hide bad spacing.
StatLab Zim is live at statlab-zim.streamlit.app and the code is at github.com/alfredshingai/statlab-zim. It still has rough edges — kaleido needs Chrome for PNG exports. But it handles an empty CSV gracefully. That's properly enough for today.
— Alfred, second coffee, probably.