Valentin Noël, PhD

Research Scientist · Mechanistic Interpretability and AI Safety · Devoteam, Paris

I study how transformer models reason, fail, and can be made safer from the inside. My work applies spectral graph signal processing to transformer attention circuits: identifying load bearing structures, detecting safety relevant failure modes, and asking whether the geometry of high dimensional representations is causally informative about model behavior.

PhD from ENS Paris Saclay and CentraleSupélec (Bayesian deep learning, inverse problems). Now at Devoteam as a Research Scientist. PI and co-PI on a Large Scale AI Research Consortium Grant (joint with a leading university and a CAC40 industrial partner). Active member of the Oxford AI Society.

Research

My core research area is mechanistic interpretability — understanding what transformer circuits actually compute, how they fail, and how failures propagate. I approach this through spectral graph signal processing: treating transformer attention matrices as graphs and analyzing their spectral properties to reveal structure invisible to activation level inspection.

Current threads: (1) spectral signatures of valid mathematical reasoning across model families; (2) the causal topology of model evolution across training checkpoints, predicting capability emergence from circuit formation; (3) scale dependent shifts in hallucination detector topology across 1B to 70B models; (4) reward modelling for grounded legal reasoning with OxAI and collaborators at Oxford.

I care about whether interpretability findings are causal and cross architecture, not merely correlational or model specific. All tools are released under open source licenses.

Selected Publications

V. Noël*, R. Franzone*, P. Wang, P. Torr, F. J. Fehr  ·  *Equal Contribution  ·  Submitted EMNLP 2026 · Under Review ICML 2026 AILaw
V. Noël, K. Healy, V. Madathil, B. Srinivasan  ·  ICML 2026 · AIWILD Workshop
V. Noël  ·  ACL 2026 · Computational Developmental Linguistics
V. Noël, T. Rodet, D. Lesselier  ·  IEEE Transactions on Computational Imaging (2024)
V. Noël, T. Rodet, D. Lesselier  ·  PIERS 2023 Best Student Paper Award

Writing

All posts →

Talks & Presentations

Projects & Demos

Run in your browser

Research projects

Spectral Glaive: Tool Call Hallucination Detection
Safety Agentic AI
Training free detection of LLM tool call hallucinations via topological collapse in transformer attention graphs. Benchmarked on the full Glaive function calling corpus across 1B to 70B models.
0.856 to 0.971 AUC · 177× speedup via Fast Fiedler GPU Lanczos · 80% recall at 90% precision
Spectral Conjecture: The Shape of Mathematical Truth
Interpretability Formal Math
Valid and invalid mathematical proofs leave distinct spectral fingerprints in transformer attention. Discovered "Memory of Refutation": historically disproven conjectures occupy a third, separate region in fingerprint space. Signal is orthogonal to perplexity — hard adversarials are more probable but spectrally anomalous.
267 labeled proofs · 60 hard adversarials · signal concentrated in late layers ("Platonic Gate")
Spectral Steering: Data Free LLM Alignment
Alignment Inference Time
Precise LLM behavior adjustment via singular value spectrum modification of MLP weights. Zero gradient updates, zero training data — operates entirely on the existing model weights at inference time. Achieves a Pareto improvement on the safety capability tradeoff.
Llama 3.2 3B: +2.4% sycophancy safety AND +1.4% GSM8K · Phi 3 Mini: 27% sycophancy reduction with +1% math
Agent Safety: Prompt Injection via Latent Probing
Safety Adversarial
Comparative study of activation probing vs spectral trust methods for prompt injection detection across GPT 2, Llama 3B, TinyLlama and Llama 1B. Probing achieves perfect in domain AUC but collapses cross domain; spectral methods generalize robustly.
Spectral: 0.86 to 0.90 AUC cross dataset · 0.757 to 0.804 F1 at 80%+ recall on mixed multi dataset
Spectral Extremism: Training Free Radicalization Detection
Safety Content Moderation
Detects radicalized text via spectral analysis of transformer attention, no fine tuning required. Discovered a critical "register confound" where models detect formality rather than ideology. Genuine within source extremism signal confirmed after register control.
0.75 to 0.82 AUROC across 5 models · Cohen's d 0.3 to 0.8 after register correction
Spectral Fingerprints: Voice Processing Across 20 Languages
Interpretability Linguistics
Training free spectral fingerprints of active vs passive voice in transformer attention across 20 typologically diverse languages. Attention heads in Layers 2 to 5 form distinct spectral signatures for syntactic structures, suggesting universal attention patterns across languages.
20 languages · cross lingual spectral consistency · submitted to ICASSP
Spectral Fold: Protein Language Model Interpretability
Bio AI Interpretability
Extends spectral graph analysis to protein language model embeddings (ESM 2, ProtT5, Ankh). Investigates whether the spectral interpretability methods developed for LLMs transfer to biological sequence models.
ESM 2, ProtT5, Ankh · Phase A: 50 proteins · scaling in progress
Spectral Prior: Geometric Tabular Data Generation
Generative Tabular
Generates synthetic tabular data with controllable spectral entropy. A 1.1M parameter NanoTabPFN, trained with spectral geometric priors, competes with commercial models 100× its size on TabArena benchmarks.
Matches or exceeds NeuralK on Blood Transfusion and Australian datasets at 1% of the parameter count

Open source

PyPI · Lead Maintainer · 2025 to Present
Open source transformer interpretability library. Real time spectral graph diagnostics for LLMs: logical fallacy detection, connectivity collapse monitoring, and circuit level trust metrics. Powers all spectral research projects above. Python and PyTorch.

Experience

Research Scientist — Devoteam
Oct 2024 to Present
Mechanistic Interpretability, AI Safety, Trustworthy AI · Paris, France
  • Spectral circuit analysis of transformer attention matrices; identified load bearing circuits and demonstrated logical coherence encoded as macroscopic topological properties (95.6% verification accuracy, d = 3.30).
  • Discovered Passive Triggered Connectivity Collapse (PTCC) as a developmental failure mode; causally localized to Layer 2 induction heads; demonstrated targeted repair via activation steering.
  • Cross architecture validation across 5+ model families (Llama, Qwen, Phi, Mistral, Gemma).
  • Agentic safety: spectral kill switches for real time contamination detection; scale dependent shifts in optimal detector topology across 1B to 70B models.
  • PI and co-PI on Large Scale AI Research Consortium Grant (joint with a leading university and CAC40 industrial partner, 2025 to Present).
PhD Researcher — ENS Paris Saclay and CentraleSupélec
Oct 2021 to Sept 2024
Deep Learning, Inverse Problems and Bayesian Methods · Supervisors: T. Rodet and D. Lesselier
  • Bayesian deep learning for uncertainty quantification in multimodal inverse problems (EM and ultrasound).
  • Transfer learning for latent structure recovery, conceptually linked to circuit analysis in foundation models.
  • HPC training on national supercomputers (SLURM); three years of university teaching in AI, Signal Processing, and Algorithms (ENS Paris Saclay and Université Paris Saclay).
Research Intern — Airbus Defence and Space and IRAP
Feb to July 2021
Hyperspectral Imaging · Toulouse, France · Supervisor: H. Carfantan
  • Forward and inverse simulation pipeline for DD CASSI hyperspectral imaging; ADMM and FISTA reconstruction; CUDA and cuPy GPU acceleration.

Education

PhD in Deep Learning — ENS Paris Saclay and CentraleSupélec
Oct 2021 to Sept 2024
Paris, France
Master in Control and Robotics (Signal Processing and AI) — Ecole Centrale de Nantes
Sept 2019 to July 2021
Nantes, France · Graduated 2nd in class

Awards

Best Presentation Award — AAAI 2026, AILaw
Association for the Advancement of Artificial Intelligence · Singapore
2026
Best Student Paper Award — PIERS
PhotonIcs and Electromagnetics Research Symposium · Prague, Czechia
2023

Technical Skills

Interpretability and Analysis
Mechanistic Interpretability, Activation Patching and Steering, Circuit Analysis, Graph Signal Processing, Spectral Methods, Causal Inference, Probing Classifiers, Attention Head Decomposition
Modeling and Training
PyTorch, JAX, Hugging Face Transformers, PEFT and LoRA, CUDA and cuPy, DeepSpeed (DDP, FSDP, ZeRO), Hugging Face Accelerate, SLURM, Bayesian Neural Networks
Evaluation and Safety
Activation Steering, Red Teaming, lm eval, Sycophancy and Faithfulness Benchmarks, Statistical Testing (bootstrap CIs, FDR correction, effect sizes)
Graph and Data
PyTorch Geometric, DGL, NetworkX, Neo4j, FAISS, Milvus, pandas, Polars
Languages
French (native), English (C2), German (B2 to C1)

Contact

I am open to conversations about mechanistic interpretability, AI safety research, and collaboration. Best reached by email (val.noel@proton.me). Also on LinkedIn and GitHub.