Skip to content

dpt model tests regress after PR #33403 #33649

Description

@dvrogozh

Folllow up from #33485 (comment). On 78b2929, there is a regression after merging this PR:

On these 2 tests:

  • tests/models/dpt/test_modeling_dpt_auto_backbone.py::DPTModelTest::test_attention_outputs
  • tests/models/dpt/test_modeling_dpt_auto_backbone.py::DPTModelTest::test_retain_grad_hidden_states_attentions

CC: @avishaiElmakies, @amyeroberts

Example output:

$ python3 -m pytest tests/models/dpt/test_modeling_dpt_auto_backbone.py::DPTModelTest::test_retain_grad_hidden_states_attentions
=================================================== test session starts ====================================================
platform linux -- Python 3.10.12, pytest-7.4.4, pluggy-1.5.0
rootdir: /home/dvrogozh/git/huggingface/transformers
configfile: pyproject.toml
plugins: pspec-0.0.4, timeout-2.3.1, hypothesis-6.112.1, xdist-3.6.1, rich-0.1.1, dash-2.18.1, cov-5.0.0, typeguard-4.3.0
collected 1 item

tests/models/dpt/test_modeling_dpt_auto_backbone.py F                                                                [100%]

========================================================= FAILURES =========================================================
__________________________________ DPTModelTest.test_retain_grad_hidden_states_attentions __________________________________

self = <tests.models.dpt.test_modeling_dpt_auto_backbone.DPTModelTest testMethod=test_retain_grad_hidden_states_attentions>

    def test_retain_grad_hidden_states_attentions(self):
        config, inputs_dict = self.model_tester.prepare_config_and_inputs_for_common()
        config.output_hidden_states = True
        config.output_attentions = self.has_attentions

        # no need to test all models as different heads yield the same functionality
        model_class = self.all_model_classes[0]
        model = model_class(config)
        model.to(torch_device)

        inputs = self._prepare_for_class(inputs_dict, model_class)

        outputs = model(**inputs)

        output = outputs[0]

        if config.is_encoder_decoder:
            # Seq2Seq models
            encoder_hidden_states = outputs.encoder_hidden_states[0]
            encoder_hidden_states.retain_grad()

            decoder_hidden_states = outputs.decoder_hidden_states[0]
            decoder_hidden_states.retain_grad()

            if self.has_attentions:
                encoder_attentions = outputs.encoder_attentions[0]
                encoder_attentions.retain_grad()

                decoder_attentions = outputs.decoder_attentions[0]
                decoder_attentions.retain_grad()

                cross_attentions = outputs.cross_attentions[0]
                cross_attentions.retain_grad()

            output.flatten()[0].backward(retain_graph=True)

            self.assertIsNotNone(encoder_hidden_states.grad)
            self.assertIsNotNone(decoder_hidden_states.grad)

            if self.has_attentions:
                self.assertIsNotNone(encoder_attentions.grad)
                self.assertIsNotNone(decoder_attentions.grad)
                self.assertIsNotNone(cross_attentions.grad)
        else:
            # Encoder-/Decoder-only models
            hidden_states = outputs.hidden_states[0]
            hidden_states.retain_grad()

            if self.has_attentions:
                attentions = outputs.attentions[0]
>               attentions.retain_grad()
E               AttributeError: 'NoneType' object has no attribute 'retain_grad'

tests/test_modeling_common.py:1677: AttributeError
===================================================== warnings summary =====================================================
src/transformers/deepspeed.py:24
  /home/dvrogozh/git/huggingface/transformers/src/transformers/deepspeed.py:24: FutureWarning: transformers.deepspeed module is deprecated and will be removed in a future version. Please import deepspeed modules directly from transformers.integrations
    warnings.warn(

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
================================================= short test summary info ==================================================
FAILED tests/models/dpt/test_modeling_dpt_auto_backbone.py::
    Here we also overwrite some of the tests of test_modeling_common.py, as DPT does not use input_ids, inputs_embeds,
    attention_mask and seq_length.
    ::test_retain_grad_hidden_states_attentions - AttributeError: 'NoneType' object has no attribute 'retain_grad'
=============================================== 1 failed, 1 warning in 2.21s ===============================================

Bisect:

$ git bisect log
git bisect start
# bad: [78b2929c0554b79e0489b451ce4ece14d265ead2] Sdpa dino v2 (#33403)
git bisect bad 78b2929c0554b79e0489b451ce4ece14d265ead2
# good: [b50ff5993a5d8b2a3d8c7558e81684f8803b044a] [`Mamba2`] Move dt calculations to kernel (#33520)
git bisect good b50ff5993a5d8b2a3d8c7558e81684f8803b044a
# good: [653eb40425344b89b5a24e7b07eb3095b04cdc9d] Add sdpa for BioGpt (#33592)
git bisect good 653eb40425344b89b5a24e7b07eb3095b04cdc9d
# good: [077b552f0780c678737700184c109066736ece41] Fix some missing tests in circleci (#33559)
git bisect good 077b552f0780c678737700184c109066736ece41
# good: [7b2b536a811c84831e2c67eb388872b7c83a8263] Fix typos (#33583)
git bisect good 7b2b536a811c84831e2c67eb388872b7c83a8263
# good: [e472e077c24d6f6f080f5535f01c48f09164ec62] Granitemoe (#33207)
git bisect good e472e077c24d6f6f080f5535f01c48f09164ec62
# good: [e71bf70e33d501810951f353f1734cb5be74b32a] Pixtral update example checkpoint (#33633)
git bisect good e71bf70e33d501810951f353f1734cb5be74b32a
# first bad commit: [78b2929c0554b79e0489b451ce4ece14d265ead2] Sdpa dino v2 (#33403)

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions