I am a machine learning and speech recognition specialist with a PhD in Automatic Speech Recognition (ASR) and ~5 years of applied research experience. My expertise lays in contextual biasing, streaming ASR, and real-world deployment in high-stakes environments (air traffic communication, call centers). I worked with low-resource speech systems, and have hands-on experience in PyTorch, Kaldi, and end-to-end ASR frameworks. Additionally, I have background in Natural Language Processing (NLP), phonetics, and linguistics.
My Research Interests:
- Audio and Speech Signal Processing
- Machine Learning and Deep Learning
- Contextual adaptation and Customisation
- ASR for assisting Air-Traffic communication
- ASR for low-resource languages (Swiss German)
- Phonetics (acoustic analysis)
You can find my CV: here (last update: July 2026)
Since a long time I have been interested in spontaneous speech phenomena, speech analysis, and speech technologies with the focus on speech recognition. Driven by a long-standing interest in signal processing and NLP, I have built a broad skill set across multiple areas of speech technologies and computational linguistics.
During my first master’s and my work at the Speech Modelling Laboratory, I spent six years working on speech perception, spontaneous speech signal, and its segmental and suprasegmental analysis. This experience motivated me to pursue further education to learn automatic language processing at the University of Zurich. For my second master’s thesis, I have trained a speech-to-text system for Swiss German dialects. Alongside my studies, I worked on a range of projects as a research assistant and intern, spanning stance detection, text classification, and sentiment analysis, as well as designing and conducting experiments on speech perception and providing technical support to phoneticians on engineering tasks.
For my PhD project, I worked on solutions for integrating context information to improve the quality of ASR predictions and the accuracy of NLP applications built on top of ASR outputs. With the focus on contextualisation and personalisation of ASR systems, my research also included streaming architectures, language modelling, and domain adaptation. To achieve my research goals, I worked with various search algorithms (Finite State Transducers, Tree-based algorithms) and different ASR frameworks (Kaldi, K2/Icefall, SpeechBrain). Working at IDIAP, I contributed to large-scale collaborative projects in high-stakes and real-world environments – (1) the European project on ASR for Air-Traffic communication and (2) research collaboration with Uniphore to improve ASR systems for call centre applications. These projects required not only the development of robust machine learning models but also careful consideration of deployment constraints such as latency, domain-specific vocabulary, and noisy input conditions. At the end of my PhD, I worked with Large Language Models, the results of which were not included in my thesis but were published at the 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW).
-
R for Data Science: https://r4ds.had.co.nz/ R cheatsheets: https://www.rstudio.com/resources/cheatsheets/ [Read More]