Skip to content

ViTNumChart: Multimodal Numerical Reasoning over Vietnamese Financial Documents

Organized by toanlnhus - Current server time: Oct. 3, 2026, 6:09 a.m. UTC

Current

Public Test
Sept. 25, 2026, midnight UTC

Next

Private Test
Oct. 9, 2026, midnight UTC

End

Competition Ends
Oct. 15, 2026, midnight UTC
Important Dates

September 18: Training Data Release

September 25: Public Test Release

October 9: Private Test Release

October 15: Result Announcement

October 25: Paper Submission

November 5: Acceptance Notification

November 12: Camera-ready Deadline

November 15: Workshop Date

Task Description

This shared task focuses on multimodal numerical reasoning over Vietnamese financial documents. Participants are required to develop systems capable of understanding and integrating information from multiple modalities, including textual content, tables, and chart images, in order to answer numerical questions about financial reports.

Unlike conventional financial question answering tasks that mainly rely on textual and tabular information, this task extends numerical reasoning to multimodal financial documents. Relevant evidence may appear in text, tables, charts, or across multiple modalities, requiring systems to identify the appropriate information and perform the necessary reasoning operations.

The expected output of a system is a reasoning program representing the sequence of operations required to solve the given question. This program-based representation makes the reasoning process explicit and enables the correctness of the reasoning procedure to be evaluated directly.

Through this shared task, we aim to promote research on multimodal financial document understanding, numerical reasoning, and interpretable reasoning systems for Vietnamese.

Task setting and participation rules
  • Single subtask: The challenge contains one official subtask.
  • Model size constraint: Each individual model used in the submitted system must have no more than 8 billion parameters (≤ 8B). Systems may consist of multiple model components or multiple stages, provided that no individual model exceeds the 8B parameter limit.
  • External data: The use of additional external data is permitted. Teams may use publicly available or self-constructed datasets in addition to the data released by the organizers.
  • Model-generated external data: If any external or additional data are generated using a generative model, the model must also satisfy the ≤ 8B parameter constraint.
  • External data disclosure: All external datasets must be clearly documented in the team's technical report.
  • External data submission: Any external data used must also be submitted or made accessible to the organizers.
  • Official ranking: Compliance with both the model-size constraint and the external-data disclosure requirements is mandatory for inclusion in the official leaderboard.
Organizers
  • Nguyen Thi Minh Huyen
  • Ha My Linh
  • Le Ngoc Toan
  • Dang Phuong Nam
  • Vu Xuan Luong
  • Pham Thi Duc
  • Ngo The Quyen
  • Phan Thi Hue
  • Le Van Cuong
Contact

Zalo Group: []

Registration

https://forms.gle/RTU4ncNog5nmis28A

References
  1. Aida Amini et al. 2019. MathQA: Towards interpretable math word problem solving with operation-based formalisms. In Proceedings of NAACL-HLT 2019.
  2. Zhiyu Chen et al. 2021. FinQA: A Dataset of Numerical Reasoning over Financial Data. In Proceedings of EMNLP 2021.
  3. Ahmed Masry et al. 2022. ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning. In Findings of ACL 2022.
Evaluation Metric

Systems are evaluated using Program Accuracy and Execution Accuracy.

Program Accuracy measures whether the predicted reasoning program is logically and mathematically equivalent to the reference reasoning program. Unlike execution-based evaluation, which considers only the final numerical answer, Program Accuracy directly evaluates the reasoning procedure used to obtain the answer.

For evaluation, the arguments of the predicted and reference programs are abstracted into symbolic variables, while the operations and their dependencies are preserved. Two programs are considered equivalent if they represent the same mathematical reasoning process, including equivalent forms resulting from commutative operations.

For example, the following programs are considered mathematically equivalent:

add(a1, a2), add(a3, a4), subtract(#0, #1)

add(a4, a3), add(a1, a2), subtract(#1, #0)

Execution Accuracy measures whether the predicted program produces the same final numerical result as the reference program when executed on the given multimodal context (text, tables, charts).

The final ranking of participating systems is determined based on their Program Accuracy on the private test set, using the official evaluation script provided by the organizers.

Evaluation Criteria

Prepare your prediction file into the following format, as a list of dictionaries, each dictionary contains two fields: the example id and the predicted program. The predicted program is a string. For example:
[
    { "qid": "bed4e850-a504-5677-9307-184caea40c71",
       "program": "add(9; 9); add(#0; 8.8); divide(#1; 3)"},
    { "qid": "e55ef634-b777-5470-9d6a-75fe742d73e3",
       "program": "subtract(76000; 61100); divide(#0; 61100); multiply(#1; 100)" },
... ]
Name the file as "results.json" and zip it. 
The final submission format is:
results.zip
•    results.json
Note that to create a valid submission, zip all the file with 'zip -r zipfilename *' starting from this directory. DO NOT zip the directory itself, just its contents. THIS IS VERY IMPORTANT.

Terms and Conditions

By participating in this shared task, you agree to the following terms:

  • Each individual model used must have no more than 8 billion parameters (≤ 8B).
  • All external datasets and data resources must be clearly documented in the team's technical report.
  • Any external data used must be submitted or made accessible to the organizers.
  • Compliance with both the model-size constraint and external-data disclosure requirements is mandatory for inclusion in the official leaderboard and final ranking.
No files have been added for this competition yet.

Public Test

Start: Sept. 25, 2026, midnight

Private Test

Start: Oct. 9, 2026, midnight

Competition Ends

Oct. 15, 2026, midnight

You must be logged in to participate in competitions.

Sign In