Skip to content

VLSP 2026 - ViSpartQA: Vietnamese Spatial Reasoning Question Answering

Organized by linhhm - Current server time: Sept. 30, 2026, 9:55 p.m. UTC

Current

Public Test
Sept. 24, 2026, 5 p.m. UTC

Next

Private Test
Oct. 8, 2026, 5 p.m. UTC

End

Competition Ends
Oct. 14, 2026, 4:59 p.m. UTC

VLSP 2026 - ViSpartQA: Vietnamese Spatial Reasoning Question Answering

ViSpartQA is a Vietnamese version of SPaRTQA, a benchmark for spatial reasoning over text. Each example has a story, which describes a scene of blocks (A, B, C) holding objects of various shapes, colors and sizes, and a set of questions about the spatial relations between those blocks and objects. Answering them takes multi-hop reasoning: transitivity, symmetry, converse relations, quantifiers ("all", "any") and negation.

Question types

  • YN: yes/no questions. The answer is Yes, No or DK (cannot be determined).
  • FB (find blocks): which blocks satisfy a condition? The answer is a possibly empty subset of the blocks.
  • FR (find relations): what are the relations between two objects? The answer is a set of relations from a fixed list: left, right, above, below, near, far, touching, DK.
  • CO (choose object): which of two described objects satisfies a relation? The options are the first, the second, both, or none.

Tracks

  • Auto track: a large machine-translated set, with about 13k training stories and 102k training questions.
  • Human track: a small set of human-written stories whose translations were checked by hand, with 37 training stories and 613 training questions.

Every team submits predictions for both tracks. Each metric is computed on each dataset separately and then averaged over the two datasets. See the Evaluation page for details.

Timeline

All deadlines are at 23:59 Vietnam time (GMT+7).

Date Event
Sep 25, 2026 Training data and public test released. The Public Test phase opens.
Oct 8, 2026 The Public Test phase ends.
Oct 9, 2026 Private test released. The Private Test phase opens.
Oct 14, 2026 System submission deadline: the Private Test phase closes.
Oct 15, 2026 Results announced.
Oct 25, 2026 Paper submission deadline.

Organizers

VLSP 2026 ViSpartQA organizing team.

Evaluation

Submission format

Each submission is one .zip file with predictions for both datasets. It must follow these rules:

  1. The zip contains exactly two files, at its root (not inside a folder):
    • auto.json: predictions for ViSpartQA-Auto.
    • human.json: predictions for ViSpartQA-Human.
    Use these names exactly, in lowercase. Names like auto_private_test.json, Auto.json or human.JSON are not accepted.
  2. The zip file itself can have any name. 
  3. Each JSON file is the test file of the current phase (auto_public_test.json / human_public_test.json, or auto_private_test.json / human_private_test.json), saved in UTF-8. Add an answer field to every question and change nothing else: keep every story and question in the same order, with the same q_id.
  4. answer is always a list, formatted by q_type as described on the Data page:
    • YN: ["Yes"], ["No"] or ["DK"].
    • FB: block names taken from candidate_answers, for example ["A", "C"]. The list can be empty ([]).
    • FR: 0-based indices into candidate_answers, for example [0, 4].
    • CO: one 0-based index into candidate_answers, for example [1].

Example layout:

MyTeam_private_test.zip
├── auto.json
└── human.json

Example question in human.json:

{"q_id": 3, "q_type": "CO", "question": "...", "candidate_answers": ["...", "...", "...", "..."], ..., "answer": [1]}

A submission fails with an error message, and gets no score, in any of these cases: a file is missing or wrongly named, a file is not valid JSON, the stories or questions do not match the test file, or a question has no answer list. An answer whose value is invalid for its type (for example ["Maybe"] for YN, or an index outside candidate_answers) is counted as wrong and listed as a warning in the detailed results.

The starting kit includes a baseline that produces a valid submission.

The private test contains additional questions that are not used for scoring. Since you cannot tell which ones they are, answer every question.

Metrics

  • Accuracy (YN and CO): a prediction is correct only when the predicted label matches the gold label. YN is a three-way classification (Yes / No / DK) and CO is a four-way classification.
    Accuracy = number of correctly answered questions / total number of questions
  • Exact Match (FR and FB, primary metric): a prediction is correct only when the predicted answer set equals the gold answer set. The order of elements does not matter.
    Exact Match = number of questions whose answer sets match exactly / total number of questions
  • Jaccard score (FR and FB, supplementary metric): the overlap between the predicted set P and the gold set G, which gives partial credit to partly correct answers.
    Jaccard(P, G) = |P ∩ G| / |P ∪ G|. If both P and G are empty, the score is 1.

Every metric is computed separately on ViSpartQA-Human and ViSpartQA-Auto. The reported result is the average of the two dataset-level scores.

Leaderboard

The leaderboard has three groups of columns, each with the same four columns:

Overall Human Auto
FR CO FB YN FR CO FB YN FR CO FB YN
  • FR and FB show Exact Match. CO and YN show Accuracy.
  • Human and Auto show the metrics of each dataset. Overall shows, for each column, the average of Human and Auto.
  • Teams are ranked by the Overall score, the mean of the four Overall columns. It equals the mean of the eight Human and Auto columns. It is listed on the detailed results page of each submission, which you can open from My Submissions.
  • The same page also has the Jaccard scores for FR and FB.

Phases

  • Public Test (Sep 25 – Oct 8, 2026): predict the answers for auto_public_test.json and human_public_test.json. You can submit up to 10 times per day.
  • Private Test (Oct 9 – Oct 14, 2026, 23:59 GMT+7): predict the answers for auto_private_test.json and human_private_test.json. You can submit up to 2 times per day and 5 times in total. The final ranking uses the leaderboard of this phase.

Every finished submission is added to the leaderboard automatically. For each team, the leaderboard shows only the submission with the highest Overall score in that phase, and it updates whenever the team submits a better one.

Terms and Conditions

  • To take part, you must sign the VLSP 2026 ViSpartQA User Agreement and send it to the organizers. List every team member on it, and register only one account per team.
  • You may use the data only for research and development of natural language processing, information retrieval or document understanding systems. Commercial use is not allowed.
  • Do not redistribute the data, in its original or modified form.
  • You may use external resources and pretrained models, including large language models. However, you must not annotate the test data by hand, and you must not use the original English SPaRTQA test answers.
  • Top teams must describe their systems in a technical report so that their results can be verified.
  • Publications that use the data must acknowledge the VLSP consortium.
  • The organizers may disqualify submissions that break these rules.

Public Test

Start: Sept. 24, 2026, 5 p.m.

Description: Public Test phase (Sep 25 - Oct 8, 2026): predict the answers of auto_public_test.json and human_public_test.json.

Private Test

Start: Oct. 8, 2026, 5 p.m.

Description: Private Test phase (Oct 9 - Oct 14, 2026): predict the answers of auto_private_test.json and human_private_test.json.

Competition Ends

Oct. 14, 2026, 4:59 p.m.

You must be logged in to participate in competitions.

Sign In