Do Transformers Parse while Predicting the Masked Word?

Zhao, Haoyu; Panigrahi, Abhishek; Ge, Rong; Arora, Sanjeev

doi:10.18653/v1/2023.emnlp-main.1029

Citation Details

Do Transformers Parse while Predicting the Masked Word?

Pre-trained language models have been shown to encode linguistic structures like parse trees in their embeddings while being trained unsupervised. Some doubts have been raised whether the models are doing parsing or only some computation weakly correlated with it. Concretely: (a) Is it possible to explicitly describe transformers with realistic embedding dimensions, number of heads, etc. that are capable of doing parsing — or even approximate parsing? (b) Why do pre-trained models capture parsing structure? This paper takes a step toward answering these questions in the context of generative modeling with PCFGs. We show that masked language models like BERT or RoBERTa of moderate sizes can approximately execute the Inside-Outside algorithm for the English PCFG (Marcus et al., 1993). We also show that the Inside-Outside algorithm is optimal for masked language modeling loss on the PCFG-generated data. We conduct probing experiments on models pre-trained on PCFG-generated data to show that this not only allows recovery of approximate parse tree, but also recovers marginal span probabilities computed by the Inside-Outside algorithm, which suggests an implicit bias of masked language modeling towards this algorithm. more »

Award ID(s):: 1845171

PAR ID:: 10530892

Author(s) / Creator(s):: Zhao, Haoyu; Panigrahi, Abhishek; Ge, Rong; Arora, Sanjeev

Publisher / Repository:: Association for Computational Linguistics

Date Published:: 2023-01-01

Page Range / eLocation ID:: 16513 to 16542

Format(s):: Medium: X

Location:: Singapore

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript
Conference Paper:
https://doi.org/10.18653/v1/2023.emnlp-main.1029

More Like this