Computer Science > Software Engineering

arXiv:2307.04346 (cs)

[Submitted on 10 Jul 2023 (v1), last revised 22 Jul 2024 (this version, v2)]

Title:Can Large Language Models Write Good Property-Based Tests?

Authors:Vasudev Vikram, Caroline Lemieux, Joshua Sunshine, Rohan Padhye

View PDF

Abstract:Property-based testing (PBT), while an established technique in the software testing research community, is still relatively underused in real-world software. Pain points in writing property-based tests include implementing diverse random input generators and thinking of meaningful properties to test. Developers, however, are more amenable to writing documentation; plenty of library API documentation is available and can be used as natural language specifications for PBTs. As large language models (LLMs) have recently shown promise in a variety of coding tasks, we investigate using modern LLMs to automatically synthesize PBTs using two prompting techniques. A key challenge is to rigorously evaluate the LLM-synthesized PBTs. We propose a methodology to do so considering several properties of the generated tests: (1) validity, (2) soundness, and (3) property coverage, a novel metric that measures the ability of the PBT to detect property violations through generation of property mutants. In our evaluation on 40 Python library API methods across three models (GPT-4, Gemini-1.5-Pro, Claude-3-Opus), we find that with the best model and prompting approach, a valid and sound PBT can be synthesized in 2.4 samples on average. We additionally find that our metric for determining soundness of a PBT is aligned with human judgment of property assertions, achieving a precision of 100% and recall of 97%. Finally, we evaluate the property coverage of LLMs across all API methods and find that the best model (GPT-4) is able to automatically synthesize correct PBTs for 21% of properties extractable from API documentation.

Subjects:	Software Engineering (cs.SE)
Cite as:	arXiv:2307.04346 [cs.SE]
	(or arXiv:2307.04346v2 [cs.SE] for this version)
	https://doi.org/10.48550/arXiv.2307.04346

Submission history

From: Vasudev Vikram [view email]
[v1] Mon, 10 Jul 2023 05:09:33 UTC (460 KB)
[v2] Mon, 22 Jul 2024 01:28:38 UTC (657 KB)

Computer Science > Software Engineering

Title:Can Large Language Models Write Good Property-Based Tests?

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Software Engineering

Title:Can Large Language Models Write Good Property-Based Tests?

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators