Syntactic Code Search with Sequence-to-Tree Matching: Supporting Syntactic Search with Incomplete Code Fragments

Matute, Gabriel (ORCID:0000000177851231); Ni, Wode (ORCID:0000000253414958); Barik, Titus (ORCID:0000000248770739); Cheung, Alvin (ORCID:0000000162616263); Chasins, Sarah_E (ORCID:0000000305573580)

doi:10.1145/3656460

Citation Details

Syntactic Code Search with Sequence-to-Tree Matching: Supporting Syntactic Search with Incomplete Code Fragments

Lightweight syntactic analysis tools like Semgrep and Comby leverage the tree structure of code, making them more expressive than string and regex search. Unlike traditional language frameworks (e.g., ESLint) that analyze codebases via explicit syntax tree manipulations, these tools use query languages that closely resemble the source language. However, state-of-the-art matching techniques for these tools require queries to be complete and parsable snippets, which makes in-progress query specifications useless. We propose a new search architecture that relies only on tokenizing (not parsing) a query. We introduce a novel language and matching algorithm to support tree-aware wildcards on this architecture by building on tree automata. We also presentstsearch, a syntactic search tool leveraging our approach. In contrast to past work, our approach supports syntactic searcheven for previously unparsable queries.We show empirically that stsea rch can support all tokenizable queries, while still providing results comparable to Semgrep for existing queries. Our work offers evidence that lightweight syntactic code search can accept in-progress specifications, potentially improving support for interactive settings. CCS Concepts: •Software and its engineering→Formal language definitions;Software maintenance tools;•Information systems→Query representation;•Theory of computation→ Tree languages. more »

Award ID(s):: 1955488 2027575

PAR ID:: 10612724

Author(s) / Creator(s):: Matute, Gabriel; Ni, Wode; Barik, Titus; Cheung, Alvin; Chasins, Sarah_E

Publisher / Repository:: Association for Computing Machinery (ACM)

Date Published:: 2024-06-20

Journal Name:: Proceedings of the ACM on Programming Languages

Volume:: 8

Issue:: PLDI

ISSN:: 2475-1421

Format(s):: Medium: X Size: p. 2051-2072

Size(s):: p. 2051-2072

Sponsoring Org:: National Science Foundation

Journal Article:
https://doi.org/10.1145/3656460

More Like this