Sketch-Driven Regular Expression Generation from Natural Language and Examples

Ye, Xi; Chen, Qiaochu; Wang, Xinyu; Dillig, Isil; Durrett, Greg

doi:10.1162/tacl_a_00339

Citation Details

Sketch-Driven Regular Expression Generation from Natural Language and Examples

Recent systems for converting natural language descriptions into regular expressions (regexes) have achieved some success, but typically deal with short, formulaic text and can only produce simple regexes. Real-world regexes are complex, hard to describe with brief sentences, and sometimes require examples to fully convey the user’s intent. We present a framework for regex synthesis in this setting where both natural language (NL) and examples are available. First, a semantic parser (either grammar-based or neural) maps the natural language description into an intermediate sketch, which is an incomplete regex containing holes to denote missing components. Then a program synthesizer searches over the regex space defined by the sketch and finds a regex that is consistent with the given string examples. Our semantic parser can be trained purely from weak supervision based on correctness of the synthesized regex, or it can leverage heuristically derived sketches. We evaluate on two prior datasets (Kushman and Barzilay 2013 ; Locascio et al. 2016 ) and a real-world dataset from Stack Overflow. Our system achieves state-of-the-art performance on the prior datasets and solves 57% of the real-world dataset, which existing neural systems completely fail on. 1 more »

Award ID(s):: 1814522

PAR ID:: 10380033

Author(s) / Creator(s):: Ye, Xi; Chen, Qiaochu; Wang, Xinyu; Dillig, Isil; Durrett, Greg

Date Published:: 2020-12-01

Journal Name:: Transactions of the Association for Computational Linguistics

Volume:: 8

ISSN:: 2307-387X

Page Range / eLocation ID:: 679 to 694

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Journal Article:
https://doi.org/10.1162/tacl_a_00339

More Like this