CVRIE
Two machine-learning studies using only scikit-learn: classifying chest X-rays, and grouping patient testimonies into symptom clusters with no labels.
- Period
- Mar 2026 – Mar 2026
- Role
- Main author of both notebooks
- Team
- 3 people
- Status
- archived
Stack and tags



A machine-learning module with a constraint that makes it interesting: only a short list of classic libraries, no deep learning and no pretrained model. Every choice has to be understood and justified. The project has two halves, each a notebook that is also the written report.
How it works
Reading chest X-rays. One image can carry several diagnoses at once, and some are far rarer than others. Two models are trained and compared honestly: a fast one, and a slower one that is a little more accurate and much less confidently wrong. Which to pick depends on whether speed or quality matters more.
Grouping patient stories. About a thousand free-text testimonies, with no categories at all. The notebook turns the words into numbers and lets a clustering algorithm find groups on its own: from three obvious ones (ear pain, burns, headaches) to fifteen readable ones, such as a mild cold or forefoot overuse pain. It is frank about the limit: most stories still fit no group.
What I took from it
- Comparing models on cost as well as accuracy gives an answer that depends on the situation, not a single winner.
- When there are no labels, reading the groups yourself is part of the evaluation.