B.Sc. B.A. Paulina Emily Hussein

© Paulina Hussein
  • Master’s Student
I am a Computer Science and NLP student with a strong research interest in the interpretability and reliability of large language models. My work focuses on mechanistic interpretability, sparse autoencoders, and the analysis of internal model representations. In several research projects, I investigated negation-related circuits in Phi-2 using sparse autoencoders and currently study copy-based in-context learning behavior in GPT-2 and Gemma-2.

My Bachelor’s thesis explored framing generation through in-context learning, examining how prompt variations influence model behavior. I also conducted a reproduction study on non-conservative force models in atomic simulations, strengthening my experience in experimental evaluation and reproducible research. Technically, I work primarily with Python, PyTorch, TransformerLens, and LogitLens. My broader goal is to contribute to trustworthy and interpretable AI systems by improving our understanding of how language models process and represent information.