ViSpartQA is a Vietnamese version of SPaRTQA, a benchmark for spatial reasoning over text. Each example has a story, which describes a scene of blocks (A, B, C) holding objects of various shapes, colors and sizes, and a set of questions about the spatial relations between those blocks and objects. Answering them takes multi-hop reasoning: transitivity, symmetry, converse relations, quantifiers ("all", "any") and negation.
Yes, No or DK (cannot be determined).Every team submits predictions for both tracks. Each metric is computed on each dataset separately and then averaged over the two datasets. See the Evaluation page for details.
All deadlines are at 23:59 Vietnam time (GMT+7).
| Date | Event |
|---|---|
| Sep 25, 2026 | Training data and public test released. The Public Test phase opens. |
| Oct 8, 2026 | The Public Test phase ends. |
| Oct 9, 2026 | Private test released. The Private Test phase opens. |
| Oct 14, 2026 | System submission deadline: the Private Test phase closes. |
| Oct 15, 2026 | Results announced. |
| Oct 25, 2026 | Paper submission deadline. |
VLSP 2026 ViSpartQA organizing team.
Each submission is one .zip file with predictions for both datasets. It must follow these rules:
auto.json: predictions for ViSpartQA-Auto.human.json: predictions for ViSpartQA-Human.auto_private_test.json, Auto.json or human.JSON are not accepted.auto_public_test.json / human_public_test.json, or auto_private_test.json / human_private_test.json), saved in UTF-8. Add an answer field to every question and change nothing else: keep every story and question in the same order, with the same q_id.answer is always a list, formatted by q_type as described on the Data page:
["Yes"], ["No"] or ["DK"].candidate_answers, for example ["A", "C"]. The list can be empty ([]).candidate_answers, for example [0, 4].candidate_answers, for example [1].Example layout:
MyTeam_private_test.zip ├── auto.json └── human.json
Example question in human.json:
{"q_id": 3, "q_type": "CO", "question": "...", "candidate_answers": ["...", "...", "...", "..."], ..., "answer": [1]}
A submission fails with an error message, and gets no score, in any of these cases: a file is missing or wrongly named, a file is not valid JSON, the stories or questions do not match the test file, or a question has no answer list. An answer whose value is invalid for its type (for example ["Maybe"] for YN, or an index outside candidate_answers) is counted as wrong and listed as a warning in the detailed results.
The starting kit includes a baseline that produces a valid submission.
The private test contains additional questions that are not used for scoring. Since you cannot tell which ones they are, answer every question.
Every metric is computed separately on ViSpartQA-Human and ViSpartQA-Auto. The reported result is the average of the two dataset-level scores.
The leaderboard has three groups of columns, each with the same four columns:
| Overall | Human | Auto | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| FR | CO | FB | YN | FR | CO | FB | YN | FR | CO | FB | YN |
auto_public_test.json and human_public_test.json. You can submit up to 10 times per day.auto_private_test.json and human_private_test.json. You can submit up to 2 times per day and 5 times in total. The final ranking uses the leaderboard of this phase.Every finished submission is added to the leaderboard automatically. For each team, the leaderboard shows only the submission with the highest Overall score in that phase, and it updates whenever the team submits a better one.
Start: Sept. 24, 2026, 5 p.m.
Description: Public Test phase (Sep 25 - Oct 8, 2026): predict the answers of auto_public_test.json and human_public_test.json.
Start: Oct. 8, 2026, 5 p.m.
Description: Private Test phase (Oct 9 - Oct 14, 2026): predict the answers of auto_private_test.json and human_private_test.json.
Oct. 14, 2026, 4:59 p.m.
You must be logged in to participate in competitions.
Sign In