Skip to content

[VisionEncoderDecoderModel] Update loss function#40863

Merged
SunMarc merged 3 commits into
huggingface:mainfrom
NielsRogge:fix_39473
Oct 14, 2025
Merged

[VisionEncoderDecoderModel] Update loss function#40863
SunMarc merged 3 commits into
huggingface:mainfrom
NielsRogge:fix_39473

Conversation

@NielsRogge

@NielsRogge NielsRogge commented Sep 13, 2025

Copy link
Copy Markdown
Collaborator

What does this PR do?

Models like Donut are currently broken on main, they can't be fine-tuned. In order to unblock users at #39473, this PR reverts #36753.

It looks like the ForCausalLMLoss class shifts the labels, however the VisionEncoderDecoderModel class does not expect shifted labels as seen here.

@HuggingFaceDocBuilderDev

Copy link
Copy Markdown

The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update.

@github-actions

github-actions Bot commented Oct 7, 2025

Copy link
Copy Markdown
Contributor

[For maintainers] Suggested jobs to run (before merge)

run-slow: vision_encoder_decoder

@NielsRogge

Copy link
Copy Markdown
Collaborator Author

cc @ArthurZucker this one can be merged

@SunMarc SunMarc left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks

@SunMarc
SunMarc merged commit 4fad35e into huggingface:main Oct 14, 2025
19 checks passed
ngazagna-qc pushed a commit to ngazagna-qc/transformers that referenced this pull request Oct 23, 2025
SangbumChoi pushed a commit to SangbumChoi/transformers that referenced this pull request Jan 23, 2026
pull Bot pushed a commit to dwongdev/transformers that referenced this pull request Jun 24, 2026
…abels[..., 1:]) (huggingface#46784)

* Fix Moonshine training-loss double-shift

Moonshine right-shifts labels into decoder_input_ids, then computes the loss via
ForCausalLMLoss, which shifts again, so it trains against labels[..., 1:]. Use a
plain CrossEntropyLoss instead, matching Whisper/Bart and the VisionEncoderDecoder
fix in huggingface#40863. Adds a regression test (with and without -100 padding).

* Regenerate Moonshine streaming model

* Mirror the training-loss regression test in moonshine_streaming

make fix-repo propagated the loss fix to the streaming model, so the same no-double-shift test should guard it there too.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants