Translation Asymmetry in LLMs as a Data Augmentation Factor: A Case Study for 6 Romansh Language Varieties
arXiv:2603.25489v2 Announce Type: replace Abstract: Recent strategies for low-resource machine translation rely on LLMs to generate synthetic data based on text in higher-resource languages.
We revisit this idea for Romansh, a language with 6 distinct varieties.
LLMs tend to confuse these varieties when translating into Romansh, but they are quite good at translating out of Romansh into a high-resource language such as German. Due to this asymmetry, the direction of data augmentation is a crucial choice.
We find that contrary to recent strategies, creating synthetic translations into the higher-resource language is the superior approach, and only this approach allows us to surpass a Gemini 3 Pro baseline on German-Romansh translation (+23 BLEU over Gemini in the lowest-resource variety).
A human evaluation confirms that our experiments yield the first model that generates fluent translations in the individual Romansh varieties.