At the moment validation bleu barely gets above zero in the tests, so they don't really prove much about our code.
we could use a larger model like sshleifer/student_marian_6_3, and more data, and train for 10 minutes . This would allows us to test whether changing default parameters/batch techniques obviously degrades performance.
The github actions CI reuses it's own disk, so this will only run there and hopefully not have super slow downloads.
At the moment validation bleu barely gets above zero in the tests, so they don't really prove much about our code.
we could use a larger model like sshleifer/student_marian_6_3, and more data, and train for 10 minutes . This would allows us to test whether changing default parameters/batch techniques obviously degrades performance.
The github actions CI reuses it's own disk, so this will only run there and hopefully not have super slow downloads.