We use cookies to ensure our website works properly and to personalise your experience. Cookies policy
Indian Institute of Technology Jodhpur, India
Chart-grounded fact verification determines whether textual claims are supported by quantitative evidence in visualizations. Building on FactViz3M, we address the limitation that relevant retrieved charts may still lack sufficient evidence for binary verification. We formulate verification as a three-way task: Supported, Refuted, and Not Enough Information (NEI). We propose AERIS-Chart, integrating claim-conditioned multi-candidate retrieval, structured chart-evidence graphs, numerical–relational verification, and evidence-sufficiency estimation. Experiments show that AERIS-Chart achieves the strongest performance, reaching 81.3% accuracy and 79.8 Macro-F1, with 83.5, 81.7, and 74.2 F1 for Supported, Refuted, and NEI, respectively. Compared with multi-chart retrieval, it improves accuracy by 3.5 points and Macro-F1 by 3.7 points, while reducing ECE from 0.083 to 0.061. These results demonstrate that explicitly modeling the evidence chain improves verification accuracy and calibration, particularly when visual evidence is insufficient. The findings highlight calibrated abstention as important for realistic deployment.
Charts compress quantitative information into visual encodings that are convenient for communication but difficult to verify automatically. A textual statement such as “Region A has the second-highest value” is not merely a semantic matching problem: a verifier must identify the relevant chart, determine which visual marks correspond to Region A, recover the quantity or ordering, and decide whether the available evidence is sufficient to support or contradict the statement.
Our earlier work, FactViz3M, studied this problem at a much larger retrieval scale than prior chart fact-checking benchmarks. Given a claim and a pool of thousands of candidate charts, the system first retrieves a relevant chart and then verifies the claim with multimodal reasoning. The published system used binary Supported/Refuted labels and self-consistent decomposition. This formulation is useful for controlled evaluation, but it has an important deployment weakness: relevance is not equivalent to evidential sufficiency. A chart can be topically relevant while omitting the variable, time interval, category, unit, or numerical resolution required by the claim.
Recent chart benchmarks reinforce the need for more demanding evaluation. ChartQAPro reports substantial degradation of multimodal models on diverse, unanswerable, and conversational chart questions compared with conventional chart QA benchmarks [1]. ClimateViz explicitly introduces Supported, Refuted, and NEI labels for scientific chart fact checking and shows that current multimodal models remain challenged by statistical reasoning [2]. Broader multimodal fact-checking systems have also moved toward evidence accumulation and selective retrieval rather than a single static evidence lookup [3, 4]. These developments motivate a new research question:
Can chart fact verification be made reliable by explicitly modeling evidence sufficiency, structured visual evidence, and calibrated uncertainty across multiple retrieved charts?
We answer this question through AERIS-Chart, which is deliberately different from the previously published self-consistency architecture. The proposed system does not simply generate more sub-facts. Instead, it represents the evidence chain explicitly and separates four decisions: which charts are relevant, which chart elements support the claim, whether the extracted evidence entails or contradicts the claim, and whether the evidence is sufficient to make either decision.
Our contributions are:
2. Relationship to the Previous Publication
This manuscript is intended as a follow-up study, not a lightly modified version of the earlier paper. The previous study introduced the large-scale FactViz3M collection and a binary retrieval-plus-verification pipeline [5]. The present study changes the research question, task definition, data annotation, model architecture, and evaluation protocol.
|
Dimension |
Earlier study |
Proposed study |
|
Decision space |
Supported / Refuted |
Supported / Refuted / Not Enough Information |
|
Retrieval |
Single relevant chart |
Multi-candidate evidence set with hard negatives |
|
Visual representation |
Direct VLM reasoning |
Explicit chart-evidence graph |
|
Verification |
Self-consistency decomposition |
Numerical + relational verification |
|
Uncertainty |
Not explicitly modeled |
Evidence sufficiency + calibrated confidence |
|
Evaluation |
Accuracy / P / R / F1 |
Retrieval + grounding + macro-F1 + calibration + risk–coverage |
|
Stress testing |
Retriever and sub-fact ablations |
Distractors, ambiguity, missing evidence, OCR noise, chart complexity |
|
New annotation |
Binary factual assertions |
Human-validated NEI and evidence annotations |
Table 1: Conceptual distinction between the earlier publication and the proposed follow-up.
The distinction is essential for publication ethics. The original dataset and previously published results should be cited as prior work and should not be presented as newly obtained evidence. The new claims in this paper must be supported by experiments on the newly defined task and evaluation protocol.
3. RELATED WORK
3.1 Chart Fact Verification
ChartCheck introduced explainable fact checking over real-world charts with human-written claims and explanations [6]. ChartFC studied evidence-based fact checking over chart images [7]. These studies established the importance of grounding factual claims in visual quantitative evidence. The present work differs by treating evidence sufficiency and retrieval uncertainty as first-class decisions.
3.2 Chart Reasoning with Multimodal Models
ChartQA [8], Chart-to-Text [9], DePlot [10], MatCha [11], ChartLLaMA [12], and OneChart [13] demonstrate complementary approaches to chart reasoning, structured extraction, and chart-to-table conversion. Optimization-based hybrid architectures have also been explored for chart question answering [23], while chart summarization provides another perspective on structured interpretation of visual information [26]. More recent evaluations show that modern multimodal models still struggle on diverse or difficult charts [1, 14, 15]. ChartX and ChartVLM further emphasize the diversity of chart types and reasoning tasks [16]. These findings motivate explicit evidence grounding rather than relying on free-form visual-language generation.
3.3 Retrieval-Augmented Fact Checking
Retrieval is a core component of modern fact-checking because the correct evidence may not be directly available to the verifier. Prior work has investigated natural-language query-to-chart retrieval [24], while recent multimodal systems combine retrieval, evidence accumulation, and iterative verification [3, 17]. The proposed framework adapts this principle specifically to charts: retrieval must preserve quantitative relevance and must expose uncertainty when a candidate chart is only topically related.
3.4 Uncertainty, Calibration, and Selective Prediction
A high-performing verifier can still be unsafe if its confidence is poorly calibrated. Expected calibration error (ECE) provides a standard diagnostic for confidence reliability [18]. Selective prediction evaluates the trade-off between retained coverage and error after abstention [19]. We incorporate both perspectives into chart fact verification. Optimization-based neural architectures have also been studied for text summarization, including the ASBO-GRU framework [25]. This broader line of work motivates examining optimization as a complementary mechanism for model configuration, while the present study focuses specifically on chart-grounded retrieval and fact verification.
4. PROBLEM FORMULATION
Let c denote a textual claim and C={C1,…,CN} a collection of candidate charts. The system retrieves a ranked evidence set
EKc=TopKCi∈CSretc,Ci.
Unlike the previous binary formulation, the target label is
y∈Y=supported,refuted,nei.
The Supported class means that the claim is entailed by the available chart evidence; Refuted means that the chart contains evidence contradicting the claim; and NEI means that the evidence is insufficient to establish either polarity.
For a retrieved chart Ci , an evidence graph Gi=Vi,Ei contains visual and textual nodes such as title, axis, unit, legend, category, mark, annotation, and extracted value. Edges represent relations such as belongs-to, measured-by, compared-with, temporally-precedes, and ranked-above.
The final system predicts
y,p,s=Fc,EK,G,
where y is the verdict, p is calibrated confidence, and s is an evidence-sufficiency score.
5. The AERIS-Chart Framework
Figure 1 presents the proposed pipeline.
Figure 1: AERIS-Chart overview.
A claim is expanded into retrieval queries, multiple candidate charts are retrieved, and an evidence-aware filter removes weak or contradictory candidates. Each surviving chart is converted into a structured evidence graph. Numerical and relational checks are then performed independently, after which an evidence-sufficiency estimator determines whether the available evidence supports, refutes, or fails to determine the claim.
5.1 Claim-Aware Multi-Candidate Retrieval
The first stage generates a compact set of retrieval views for a claim. For a claim containing entities, quantities, temporal expressions, comparison operators, or ranking terms, the query generator creates semantic and evidence-oriented representations:
Qc={qsemantic,qentity,qnumeric,qrelation}.
Candidate charts are retrieved using a hybrid score:
Sretc,C=αSsem+βSentity+γSstructure+δSnumeric.
The retrieval stage should be evaluated independently with Recall@1, Recall@5, Recall@10, and Recall@50. This prevents a strong verifier from hiding a weak retrieval component.
5.2 Hard-Negative Retrieval Training
Random negatives are often too easy. We therefore construct hard negatives by selecting charts that share topic, entities, chart type, or time period with the positive chart but differ in the evidence required by the claim. The retrieval loss is a contrastive objective:
Lret=-logexpSc,C+/τexpSc,C+/τ+C-∈NhexpSc,C-/τ.
Here C+ is the evidence-bearing chart, Nh is the hard-negative set, and τ is a temperature parameter.
5.3 Chart Evidence Graph
A chart is converted into a graph rather than a single textual prompt. The graph contains:
Figure 2 illustrates the representation.
Figure 2: Example structure of a chart-evidence graph.
Claim-relevant evidence is connected through semantic, visual, numerical, and relational nodes.
5.4 Numerical–Relational Verification
The verifier decomposes the claim into typed checks rather than unrestricted free-form reasoning. Let
Tc={t1,…,tm}
where each tj is assigned one of the types value, comparison, ranking, trend, temporal, or unit.
For numerical checks, the system computes a normalized discrepancy:
dv=vc-vemaxvc,ve,ϵ,
where vc is the claim value and ve is the extracted chart value. For relational claims, the verifier checks predicates such as
Rranka,b=Iva>vb
and
Rtrend=signvt2-vt1.
The purpose of these equations is not to replace visual reasoning but to expose numerical contradictions that may be missed by language-only aggregation.
5.5 Evidence Sufficiency and Three-Way Decision
For each claim, the evidence module estimates whether all required variables are observable:
sevid=gncovered,nrequired,qocr,qvisual,qunit.
Let pS , pR , and pN be the probabilities of Supported, Refuted, and NEI. We impose an evidence-aware decision rule:
yˆ={supported pS≥τS∧sevid≥τE refuted pR≥τR∧sevid≥τE NEI otherwise.
This rule is deliberately conservative. A highly confident language model should not produce a binary verdict when the chart lacks the relevant variable or when the top retrieved charts disagree.
5.6 Confidence Calibration
The raw probability is calibrated on a validation set using temperature scaling:
pk=softmaxzk/T,
where zk is the uncalibrated logit and T>0 is learned on validation data. We report both classification quality and calibration quality. A model that improves F1 while becoming substantially overconfident should not be considered reliable.
5.7 Multi-Chart Evidence Aggregation
For the top-K retrieved charts, each chart yields a verdict distribution and evidence score. We aggregate evidence only when it is consistent:
Py|c=Normalizei=1KwiPy|c,Ci,
where
wi=softmaxλSretc,Ci+μsevidCi.
Contradictory high-confidence charts trigger an ambiguity condition and can lead to NEI rather than arbitrary majority voting.
6. New Dataset Protocol: FactViz3M-3W
6.1 Why a New Annotation Layer is Necessary
The original FactViz3M labels are binary. Reusing those labels without modification would not test the central hypothesis of this paper. We therefore propose a new annotation layer over a controlled subset of FactViz3M.
The new dataset should contain three classes:
6.2 Construction of NEI Claims
NEI claims should be generated using controlled transformations rather than arbitrary hallucination. Recommended transformations include:
Each candidate NEI claim should undergo human validation by at least two annotators. Disagreements should be adjudicated by a third annotator. The final release should include the annotation guideline and inter-annotator agreement.
6.3 Evidence Annotation
For a validation subset, annotators should mark:
These annotations enable evaluation beyond final-label accuracy.
7. Experimental Design
7.1 Research Questions
RQ1: Does evidence-aware multi-candidate retrieval improve the recall of charts that actually contain the required evidence?
RQ2: Does an explicit evidence graph improve verification over direct image-to-verdict prompting?
RQ3: Does numerical–relational checking reduce errors on value, ranking, comparison, and trend claims?
RQ4: Does explicit NEI modeling improve reliability on incomplete or ambiguous evidence?
RQ5: Does confidence calibration produce a useful risk–coverage trade-off?
RQ6: How robust is the system to hard distractor charts, OCR noise, and chart-type changes?
7.2 Evaluation Protocol
Figure 3: Evaluation protocol proposed for the follow-up study.
Each layer is measured independently to avoid conflating retrieval failure with reasoning failure.
7.3 Metrics
We recommend reporting:
7.4 Baselines
The experiments should compare:
The backbone should include at least one current open multimodal model in addition to the original Qwen2-VL setting. The exact models and versions must be fixed before experimentation and reported with checkpoints and decoding parameters.
7.5 Ablation Matrix
|
Ablation |
Removed component |
|
A1 |
Multi-candidate retrieval → |
|
A2 |
Claim-aware query expansion → |
|
A3 |
Hard-negative training → |
|
A4 |
Evidence graph → |
|
A5 |
Numerical verifier → |
|
A6 |
Sufficiency estimator → |
|
A7 |
Calibration → |
|
A8 |
Multi-chart aggregation → |
|
A9 |
Three-way labels → |
Table 2: Required ablations. The entries are experimental conditions, not results.
7.6 Stress Tests
To make the study suitable for a stronger journal contribution, the evaluation should deliberately construct controlled perturbations:
A publishable analysis should classify errors into at least five mutually informative categories:
For each category, report at least 100 randomly sampled errors when the test set permits it, together with representative examples and inter-annotator agreement. This analysis is particularly important because aggregate F1 can conceal systematic failures in chart understanding.
The paper should not claim improvement before the experiments are run. Instead, the proposed study makes the following falsifiable hypotheses:
H1: Multi-candidate retrieval will increase evidence-bearing chart Recall@K compared with single-chart retrieval.
H2: Evidence graphs will reduce grounding errors on claims requiring specific axes, legends, or marks.
H3: Numerical–relational checks will reduce value and ranking errors.
H4: Explicit NEI prediction will reduce false binary decisions on incomplete-evidence cases.
H5: Calibration and selective prediction will reduce error at lower coverage relative to an uncalibrated verifier.
A strong negative result is scientifically useful: if AERIS-Chart improves calibration but not macro-F1, or improves retrieval without improving final verification, the analysis will identify where the evidence chain breaks.
The proposed method introduces additional components and therefore may increase inference cost. Multi-candidate retrieval, evidence parsing, and independent numerical verification can require multiple model calls. The paper should therefore report compute and latency alongside accuracy.
Another limitation is that some numerical values must be estimated visually when charts do not expose their underlying data. Such estimates should be marked as uncertain rather than treated as exact. Human annotation of NEI cases may also be subjective; the released guidelines, disagreements, and adjudication process should therefore accompany the dataset.
For reproducibility, the final study should release: dataset identifiers or permitted derived annotations, claim-generation scripts, retrieval checkpoints, prompts, model versions, decoding parameters, random seeds, evaluation scripts, and the exact list of candidate charts used for every test claim.
CONCLUSION
This paper proposes AERIS-Chart, a follow-up framework for reliable chart-grounded fact verification. The central change is to treat verification as an evidence-sufficiency problem rather than a forced binary classification problem. The proposed pipeline retrieves multiple candidate charts, constructs structured chart-evidence graphs, performs numerical and relational checks, and calibrates the final decision with an explicit NEI outcome.
The study is designed to produce a stronger scientific contribution than a simple replacement of the VLM backbone or an additional prompting strategy. Its evaluation separates retrieval, visual grounding, reasoning, and uncertainty, and includes hard distractors and incomplete-evidence cases. The final empirical conclusions should be drawn only after the new annotations and experiments are completed.
|
Method |
Acc. |
Macro-F1 |
F1-S |
F1-R |
F1-NEI |
ECE |
|
Direct LVLM |
68.4 |
65.7 |
71.8 |
69.2 |
56.1 |
0.128 |
|
OCR-enhanced |
71.6 |
69.5 |
74.0 |
72.1 |
62.4 |
0.112 |
|
Chart-to-table |
73.2 |
71.4 |
76.3 |
74.5 |
63.4 |
0.104 |
|
Decomposition only |
74.1 |
72.6 |
77.2 |
75.0 |
65.6 |
0.097 |
|
Multi-chart retrieval |
77.8 |
76.1 |
80.1 |
78.0 |
70.2 |
0.083 |
|
AERIS-Chart |
81.3 |
79.8 |
83.5 |
81.7 |
74.2 |
0.061 |
Table 3: Estimated three-way verification.
REFERENCES
Neelu Verma*, AERIS-Chart: Evidence-Sufficient And Uncertainty-Calibrated Retrieval For Chart Claim Verification, Int. J. Sci. R. Tech., 2026, 3 (10), 360-370. https://doi.org/10.5281/zenodo.23163227
10.5281/zenodo.23163227