How far can we get with one GPU in 100 hours? CoAStaL at MultiIndicMT Shared Task

Research output: Chapter in Book/Report/Conference proceedingArticle in proceedingsResearchpeer-review

Documents

  • Fulltext

    Final published version, 327 KB, PDF document

This work shows that competitive translation results can be obtained in a constrained setting by incorporating the latest advances in memory and compute optimization. We train and evaluate large multilingual translation models using a single GPU for a maximum of 100 hours and get within 4-5 BLEU points of the top submission on the leaderboard. We also benchmark standard baselines on the PMI corpus and re-discover well-known shortcomings of translation systems and metrics.
Original languageEnglish
Title of host publicationProceedings of the 8th Workshop on Asian Translation (WAT2021)
PublisherAssociation for Computational Linguistics
Publication date2021
Pages205-211
DOIs
Publication statusPublished - 2021
Event8th Workshop on Asian Translation (WAT2021) - Online
Duration: 5 Aug 20216 Aug 2021

Conference

Conference8th Workshop on Asian Translation (WAT2021)
ByOnline
Periode05/08/202106/08/2021

ID: 300450019