Co-Evolution for Logic: Improving Autoformalization with Large Language Models
Mark O'Neill ⋅ Megan Ma
Abstract
Autoformalisation (translating natural language into formal logic) has applications in law, mathematics, and AI. While large language models show promise, they often produce incorrect or syntactically invalid outputs. We adapt a co-evolutionary framework (CoCoEvo) to an autoformalisation task and extend it with syntax repair, diversity mechanisms, and validation tests (CoCoEvo++). On the FOLIO benchmark, CoCoEvo++ achieves $58.4 \pm 3.9\%$ correctness, up from $39.3 \pm 3.5\%$ for single-pass generation. On the SARA statutory reasoning benchmark, CoCoEvo++ achieves $57.7 \pm 2.2\%$, up from $49.3 \pm 1.9\%$ for single-pass, with more moderate gains indicating the limits of single-objective co-evolution on multi-objective legal reasoning tasks. We highlight limitations, such as reliance on a fixed vocabulary and increased computational cost, and point to directions for future research in more adaptive and efficient methods.
Chat is not available.
Successful Page Load