Valentin Noël, PhD
Research Scientist · Mechanistic Interpretability and AI Safety · Devoteam, Paris
I study how transformer models reason, fail, and can be made safer from the inside. My work applies spectral graph signal processing to transformer attention circuits: identifying load bearing structures, detecting safety relevant failure modes, and asking whether the geometry of high dimensional representations is causally informative about model behavior.
PhD from ENS Paris Saclay and CentraleSupélec (Bayesian deep learning, inverse problems). Now at Devoteam as a Research Scientist. PI and co-PI on a Large Scale AI Research Consortium Grant (joint with a leading university and a CAC40 industrial partner). Active member of the Oxford AI Society.
Research
My core research area is mechanistic interpretability — understanding what transformer circuits actually compute, how they fail, and how failures propagate. I approach this through spectral graph signal processing: treating transformer attention matrices as graphs and analyzing their spectral properties to reveal structure invisible to activation level inspection.
Current threads: (1) spectral signatures of valid mathematical reasoning across model families; (2) the causal topology of model evolution across training checkpoints, predicting capability emergence from circuit formation; (3) scale dependent shifts in hallucination detector topology across 1B to 70B models; (4) reward modelling for grounded legal reasoning with OxAI and collaborators at Oxford.
I care about whether interpretability findings are causal and cross architecture, not merely correlational or model specific. All tools are released under open source licenses.
Selected Publications
Full list on Google Scholar.
Writing
All posts →Talks & Presentations
Projects & Demos
Run in your browser
Research projects
Open source
Experience
- Spectral circuit analysis of transformer attention matrices; identified load bearing circuits and demonstrated logical coherence encoded as macroscopic topological properties (95.6% verification accuracy, d = 3.30).
- Discovered Passive Triggered Connectivity Collapse (PTCC) as a developmental failure mode; causally localized to Layer 2 induction heads; demonstrated targeted repair via activation steering.
- Cross architecture validation across 5+ model families (Llama, Qwen, Phi, Mistral, Gemma).
- Agentic safety: spectral kill switches for real time contamination detection; scale dependent shifts in optimal detector topology across 1B to 70B models.
- PI and co-PI on Large Scale AI Research Consortium Grant (joint with a leading university and CAC40 industrial partner, 2025 to Present).
- Bayesian deep learning for uncertainty quantification in multimodal inverse problems (EM and ultrasound).
- Transfer learning for latent structure recovery, conceptually linked to circuit analysis in foundation models.
- HPC training on national supercomputers (SLURM); three years of university teaching in AI, Signal Processing, and Algorithms (ENS Paris Saclay and Université Paris Saclay).
- Forward and inverse simulation pipeline for DD CASSI hyperspectral imaging; ADMM and FISTA reconstruction; CUDA and cuPy GPU acceleration.
Education
Awards
Technical Skills
Contact
I am open to conversations about mechanistic interpretability, AI safety research, and collaboration. Best reached by email (val.noel@proton.me). Also on LinkedIn and GitHub.