September 18: Training Data Release
September 25: Public Test Release
October 9: Private Test Release
October 15: Result Announcement
October 25: Paper Submission
November 5: Acceptance Notification
November 12: Camera-ready Deadline
November 15: Workshop Date
This shared task focuses on multimodal numerical reasoning over Vietnamese financial documents. Participants are required to develop systems capable of understanding and integrating information from multiple modalities, including textual content, tables, and chart images, in order to answer numerical questions about financial reports.
Unlike conventional financial question answering tasks that mainly rely on textual and tabular information, this task extends numerical reasoning to multimodal financial documents. Relevant evidence may appear in text, tables, charts, or across multiple modalities, requiring systems to identify the appropriate information and perform the necessary reasoning operations.
The expected output of a system is a reasoning program representing the sequence of operations required to solve the given question. This program-based representation makes the reasoning process explicit and enables the correctness of the reasoning procedure to be evaluated directly.
Through this shared task, we aim to promote research on multimodal financial document understanding, numerical reasoning, and interpretable reasoning systems for Vietnamese.
Zalo Group: []
https://forms.gle/RTU4ncNog5nmis28A
Systems are evaluated using Program Accuracy and Execution Accuracy.
Program Accuracy measures whether the predicted reasoning program is logically and mathematically equivalent to the reference reasoning program. Unlike execution-based evaluation, which considers only the final numerical answer, Program Accuracy directly evaluates the reasoning procedure used to obtain the answer.
For evaluation, the arguments of the predicted and reference programs are abstracted into symbolic variables, while the operations and their dependencies are preserved. Two programs are considered equivalent if they represent the same mathematical reasoning process, including equivalent forms resulting from commutative operations.
For example, the following programs are considered mathematically equivalent:
add(a1, a2), add(a3, a4), subtract(#0, #1)
add(a4, a3), add(a1, a2), subtract(#1, #0)
Execution Accuracy measures whether the predicted program produces the same final numerical result as the reference program when executed on the given multimodal context (text, tables, charts).
The final ranking of participating systems is determined based on their Program Accuracy on the private test set, using the official evaluation script provided by the organizers.
Prepare your prediction file into the following format, as a list of dictionaries, each dictionary contains two fields: the example id and the predicted program. The predicted program is a string. For example:[ { "qid": "bed4e850-a504-5677-9307-184caea40c71", "program": "add(9; 9); add(#0; 8.8); divide(#1; 3)"}, { "qid": "e55ef634-b777-5470-9d6a-75fe742d73e3", "program": "subtract(76000; 61100); divide(#0; 61100); multiply(#1; 100)" },... ]
Name the file as "results.json" and zip it.
The final submission format is:
results.zip
• results.json
Note that to create a valid submission, zip all the file with 'zip -r zipfilename *' starting from this directory. DO NOT zip the directory itself, just its contents. THIS IS VERY IMPORTANT.
Terms and Conditions
By participating in this shared task, you agree to the following terms:
Start: Sept. 25, 2026, midnight
Start: Oct. 9, 2026, midnight
Oct. 15, 2026, midnight
You must be logged in to participate in competitions.
Sign In