fix failed test cases for qwen3_omni_moe model#47449
Conversation
Signed-off-by: kaixuanliu <[email protected]>
Signed-off-by: kaixuanliu <[email protected]>
…on_ids_and_fa_kwargs` Signed-off-by: kaixuanliu <[email protected]>
Signed-off-by: kaixuanliu <[email protected]>
| _no_split_modules = [ | ||
| "Qwen3OmniMoeThinkerTextDecoderLayer", | ||
| "Qwen3OmniMoeVisionBlock", | ||
| ] | ||
|
|
There was a problem hiding this comment.
could you please get all no-split modules and define them once in here? When loading a backbone, we'll filter all available modules from the list anyway and it makes it easy if we keep one list
I am seeing that vision model override it and define ["Qwen3OmniMoeVisionEncoder"], then audio overrides as ["Qwen3OmniMoeAudioEncoder"] atm
There was a problem hiding this comment.
Sure, done for it.
Signed-off-by: kaixuanliu <[email protected]>
|
[For maintainers] Suggested jobs to run (before merge) run-slow: qwen3_omni_moe |
CI recapDashboard: View test results in Grafana |
|
The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update. |
* upstream/main: (39 commits) Remove deprecated training args and `is_fast` property (huggingface#46917) Consistent output shape from `get_image_features` (huggingface#46405) Fix multi-device mxfp4 dequantization race in `_convert_moe_packed_tensors` (huggingface#47423) fix failed test cases for qwen3_omni_moe model (huggingface#47449) Fix Hunyuan-VL PIL image resize parity with reference preprocessing (huggingface#47233) Move `value` padding into the attention interfaces that need it (huggingface#47451) Simplify function dispatch for linear attention (huggingface#47450) [cache] Allow sliding window layers to be roll-backed for speculative decoding (huggingface#47447) Fix double-shifted training loss in GitForCausalLM (huggingface#47395) Fix CohereASR training-loss double-shift (same as Moonshine fix huggingface#46784) (huggingface#46895) Warn when `group_by_length` is silently ignored for iterable datasets (huggingface#47379) Update bug report list (huggingface#46607) Fix shape mismatch in KyutaiSpeechToText `generate()` last window (huggingface#46952) Optimize flash attention max seqlen computation in vision attention (huggingface#47170) fix: remove unreachable return in special token builder (huggingface#47420) Add Harry to slow CI (huggingface#47454) BLT: vectorize patch length processing (huggingface#47385) Fix `TrackioCallback` fails to log evaluation metrics after training ends (huggingface#46935) [Kimi] add integration tests (huggingface#47383) Fix typo in `MusicgenForCausalLM.generate()` (huggingface#46974) ...
This PR tries to fix following failed test cases:
@zucchini-nlp @ydshieh pls help review, thx!