Mammogram quality assessment
MSc Computer Science · In progress
Automated PGMI scale classification using a ConvNeXt ensemble to assess mammogram image quality with deep learning.
Co-founder & machine learning engineer
My research interests are LLM training, evaluation, and reliable reasoning. My current public experiments study how language models reason and how to evaluate their behavior. Earlier, I worked on deep learning for medical image analysis at the University of Innsbruck.
Independent experiment · July 2026
I evaluated Qwen2.5-7B-Instruct on a bookkeeping task with reasoning depths of 4, 8, 16, and 32, using 32 tasks per condition and depth. The experiment compares chain-of-thought prompting with an external state ledger and ablations of its checking mechanism.
At depth 32, the ledger condition reached 50.0% accuracy versus 28.1% for the chain-of-thought baseline. Removing the checking mechanism produced the same accuracy at every tested depth: the checks added no measured accuracy in Stage A.
These preliminary results cover a single model and task family. This is an inference-time evaluation using a fixed pretrained model. The report includes the experimental setup, confidence intervals, ablations, and limitations; code and recorded results are public.
MSc Computer Science · In progress
Automated PGMI scale classification using a ConvNeXt ensemble to assess mammogram image quality with deep learning.
BSc Computer Science · 2017–2020
Automatic semantic segmentation of kidneys and kidney tumors in medical images, combining YOLO-based localization with U-Net segmentation.
At Heliotherm, I applied machine learning to predictive maintenance for more than 4,000 heat pumps. At Accemic, I contributed to the EU-funded COEMS and TRISTAN research programs.