DOI: 10.3390/batteries12080314 ISSN: 2313-0105

Reconciling Manufacturer Claims with Measured Degradation in Commercial Lithium-Ion Cells: A Provenance-Aware Knowledge Graph with Coverage-Gated Abstention

Alexandru Lecu, Lezan Hawizy, Adrian Groza

Manufacturer datasheets state battery cycle life under conditions that rarely match how cells are used, while public cycling datasets measure degradation under conditions datasheets do not cover. We present a knowledge-graph (KG) system that represents claims, measurements, and independent tests of commercial lithium-ion cells with full provenance, detects claim-versus-measured and claim-versus-claim discrepancies conditioned on the comparability of test conditions, and supports cycle-life prediction with coverage-gated abstention. On the 124-cell Severson dataset under leave-one-policy-group-out cross-validation, graph-derived neighbor features do not significantly improve point prediction over a strong early-cycle baseline (RMSE 135 vs. 141 cycles), but graph coverage provides a statistically significant abstention signal (Spearman ρ=−0.25 with prediction error, p=0.006) that reduces retained RMSE by roughly 40% at 60% retention, where random abstention does not. Deployed zero-shot on a second cycling study of the same commercial cell, the gate abstained on all 77 cells; the counterfactual confirms every refusal (approximately 83% error had it answered), an error an ungated baseline commits silently. On a third study with commensurable features, the gate’s first partial acceptance (17 of 45 cells) is itself diagnostic: coverage acts partly as a lifetime proxy out of distribution, and five labeled cells halve retained error while leaving that proxy in place—adaptation repairs the predictor, not the selection criterion. A 70B open-weight LLM extracts datasheet claims at F1=0.70 with non-deterministic output even at temperature 0; a deterministic validator with three-run consensus raises this to F1=0.78 with zero unsourced values; on a held-out datasheet, precision and the zero-unsourced-value property transfer while recall falls to 0.34, localizing the extractor’s boundary at table-structured content; row-level table grounding, implemented in response, raises held-out recall to 0.63 with zero hallucinations at a measured precision cost. Reconciling claims across document variants shows that roughly one in three cross-document specification comparisons (14 of 43, three commercial cells) yields a conflict or condition mismatch, twelve involving third-party documents and two internal to a single manufacturer’s own documents. A hand-labeled, condition-annotated gold standard of 103 claims (62 development, 41 held-out; inter-annotator κ=0.74 on property naming) and a staged, human-gated literature-monitoring pipeline are released with the code.

More from our Archive