Iñigo Parra

Portrait of Iñigo Parra

San Francisco, CA

About Me

I am a PhD student in Linguistics at UC Berkeley working on interpretability and on how deep learning can help us understand language and human cognition. I am also currently a research scientist intern at Snowflake, where I work on reinforcement learning, agent behavioral analysis, and model routing.

Latest Publications

Method overview: language-model embeddings predict brain responses through ridge regression, and the geometry of both spaces is compared with Gromov–Wasserstein distance and CKA.

Neural Correlates of Language Models Are Specific to Human Language

NeurIPS 2025 / PMLR

TL;DR: This study investigates the neural correlates between language models and human brain activity, demonstrating that the language representations underlying modern language models align specifically with human language processing patterns rather than general sequence processing. I also show shared geometrical properties of LM and brain representations.

Paper
UMAP projection of sparse autoencoder codes, colored by k-means cluster.

Interpretable Sparse Features for Probing Self-Supervised Speech Models

AACL 2025

TL;DR: I proposed a novel application of SAEs to understand self-supervised speech models through interpretable sparse feature extraction. The method revealed linguistically meaningful patterns in SAE features, offering new insights into what these models learn about phonetic and phonological structure.

Paper
Architecture diagram: a dynamic threshold selection module decides how the input is routed through a stack of layers.

Adaptive Compute Efficient Learning via Conceptual-Criticality

AAAI 2025

TL;DR: We introduced a conceptual-criticality framework for adaptive compute allocation in neural networks, enabling models to dynamically adjust computational resources based on input complexity. This approach achieves significant efficiency gains while maintaining or improving model performance.

Paper
Loss and perplexity curves over training steps for English, Spanish, French, Basque, Hungarian, and Finnish.

Morphological Typology in BPE Subword Productivity and Language Modeling

NeurIPS 2024

TL;DR: This project explores how morphological typology affects BPE tokenization efficiency and language model performance across diverse languages. I demonstrated that morphology has a major effect on tokenization, and that this also affects training.

Paper
Token processing diagram: each token is copied, one copy is degraded and then enhanced, and the final dataset keeps both the original and the enhanced versions.

Noise Be Gone: Does Speech Enhancement Distort Linguistic Nuances?

ACL 2024

TL;DR: I investigated whether speech enhancement techniques preserve or distort fine-grained linguistic features. My findings revealed that noise reduction improves overall intelligibility, and that enhancement methods do not significantly alter phonetic distinctions critical for downstream linguistic analysis.

Paper
Summary of the approach: masked pronoun and linguistic-unit queries are given to multilingual and monolingual BERT and DistilBERT models, followed by quantitative and qualitative analysis.

UnMASKed: Quantifying Gender Biases in Language Models through Linguistically Informed Job Market Queries

EACL 2024

TL;DR: Using linguistically informed job market queries, I developed a methodology to quantify gender biases in masked language models. The analysis revealed systematic biases in occupational associations and provides insights into how language models perpetuate societal stereotypes.

Paper

Visuals

Sinusoidal Position Encodings

Demonstrating how sinusoidal positional encodings work in transformer models

Query, Key, and Value

Visualizing the Query, Key, and Value matrices in attention mechanisms

The GAN Objective

Breaking down the mathematical formulation of Generative Adversarial Networks

Some of my favourite quotes

“All models are wrong, but some are useful.”

George E. P. Box (1919–2013)

“Intuitions have been tacitly granted a privileged position in generative grammar. The result has been the construction of elaborate theoretical edifices supported by disturbingly shaky empirical evidence.”

Wasow & Arnold (2004, p. 1482)

“Being a native speaker doesn’t confer papal infallibility on one’s intuitive judgments.”

Raven McDavid (1985)