Machine Translation (MT) in the legal domain is challenging because legal documents require high accuracy and consistent use of specialized terminology. Legal texts often contain complex sentences, domain-specific terms, and expressions whose meaning depends strongly on the legal context. Inconsistent translation of the same legal term can lead to ambiguity or changes in meaning. These challenges become more difficult when using relatively small models with limited computational resources.
The domain is legal, and the supported languages are English and Vietnamese. Systems are evaluated in both directions:
en-vi)vi-en)All teams receive the same legal-domain training, development, and test data. The restricted-resource setting encourages effective domain adaptation, terminology-aware methods, and efficient training rather than reliance on larger models alone. See the Terms and Conditions for model, data, and inference requirements.
Submissions are evaluated using corpus-level SacreBLEU-style BLEU (13a tokenization, exponential smoothing) and chrF (character order 6, beta 2). For a phase containing both language directions, each metric is calculated separately for en-vi and vi-en, then macro-averaged so both directions contribute equally. Both metrics range from 0 to 100, and higher scores are better.
Upload one ZIP archive containing a file named exactly results.csv at the root of the archive:
submission.zip โโโ results.csv
Do not place results.csv inside another directory.
The file must be UTF-8 encoded and contain exactly these columns:
sample_id: the unchanged identifier supplied with the test source;direction: either en-vi or vi-en;translation: the complete translated document. The alias prediction is also accepted.Example:
sample_id,direction,translation Bat_dong_san/137,en-vi,"<complete Vietnamese translation>" Bat_dong_san/137,vi-en,"<complete English translation>"
Translations may contain commas, quotation marks, and line breaks. Produce a standards-compliant CSV: quote such fields and escape an internal quotation mark by doubling it.
sample_id, for exactly 100 prediction rows.Rows may appear in any order. Every required ID/direction combination must appear exactly once.
sample_id or the provided direction.A submission is rejected if the ZIP does not contain results.csv, required columns are missing, a direction is invalid, a prediction is empty, an ID/direction is duplicated, or the submitted rows do not match the phase test set.
All submissions must comply with the model, data, and inference requirements in the Terms and Conditions.
By participating in the VLSP 2026 Legal Machine Translation shared task, a team agrees to the following conditions.
Start: Sept. 25, 2026, midnight
Start: Oct. 9, 2026, midnight
Oct. 15, 2026, midnight
You must be logged in to participate in competitions.
Sign In