GHA-Agent: A Multi-Agent Framework for GitHub Actions Log Parsing
Chuyue Wu, Mingyuan Zhang, Yinggang Ling, Pinjia He, Sijie Xu, Yao Li, Tao ZhangThe advent of continuous integration and continuous deployment (CI/CD) technologies has significantly reduced development cycles for engineering teams. As the world's largest open-source repository, the GitHub Actions (GHA) platform has rapidly become the preferred tool for developers. However, GHA failures generate massive volumes of unstructured logs, making root cause analysis (RCA) exceptionally challenging. Existing log parsing research primarily focuses on structured system logs, relying on static templates that struggle to handle the highly dynamic and unstructured nature of GHA logs. Although recent explorations into unstructured log parsing have emerged, two core challenges persist in the GHA context: First, there is a lack of a standard benchmark dataset covering the entire process from log parsing to code repair. Second, it is difficult to semantically distinguish among different failure modes such as build errors, test failures, and runtime exceptions. This study proposes a multi-agent log parsing framework and the end-to-end evaluation benchmark GHA-Bench to address these challenges. GHA-Agent is built on an LLM-driven multi-agent architecture for zero-shot error localization in extremely long logs. It integrates autonomous tool invocation with context-aware parsing strategy selection to enable effective reasoning over highly unstructured log data. GHA-Bench includes failure logs from 323 real repositories with manually labeled parsing results and corresponding golden patches, supporting comprehensive evaluation. Experimental results demonstrate that GHA-Agent achieves 97.75% error coverage in log parsing. In a downstream manual evaluation, experts accepted the generated log diagnoses for 83.0% of instances and the root-cause reports for 78.0% of instances, indicating that GHA-Agent often preserves evidence useful for code-level diagnosis.