Fix QwenImage txt_seq_lens handling #12702

kashif · 2025-11-23T18:03:55Z

What does this PR do?

Removes the redundant txt_seq_lens plumbing from all QwenImage pipelines and modular steps; the transformer now infers text length from encoder inputs/masks and validates optional overrides.
Builds a lightweight broadcastable attention mask from encoder_hidden_states_mask inside the double-stream attention, avoiding full seq_len² masks while keeping padding tokens masked.
Adjusts QwenImage Transformer/ControlNet RoPE to take a single text length and documents the fallback behavior.
Adds regression tests to ensure short txt_seq_lens values and encoder masks are handled safely.

Before submitting

This PR fixes a typo or improves the docs (you can dismiss the other checks if that's the case).
Did you read the contributor guideline?
Did you read our philosophy doc (important for complex PRs)?
Was this discussed/approved via a GitHub issue or the forum? Please add a link to it if that's the case.
Did you make sure to update the documentation with your changes? Here are the
documentation guidelines, and
here are tips on formatting docstrings.
Did you write any new necessary tests?

Who can review?

Anyone in the community is free to review the PR once the tests have passed. Feel free to tag
members/contributors who may be interested in your PR.

HuggingFaceDocBuilderDev · 2025-11-23T18:12:12Z

The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update.

dxqb · 2025-11-29T06:54:47Z

just a few comments, not a full review:

there is some overlap with Fix qwen encoder hidden states mask #12655
this code has the same issue mentioned in Fix qwen encoder hidden states mask #12655 (expecting boolean semantics in a FloatTensor - but float attention masks are interpreted differently)
Could you clarify what the purpose of this PR is?
If the purpose is to remove the txt_seq_lens parameters, and infer the sequence lengths from the attention mask: why is it still a parameter of the transformer model?
If the purpose is towards passing sequence lengths to the attention dispatch (see Qwen Image: txt_seq_lens is redundant and not used #12344 (comment)), the sequence lengths for each batch sample must be inferred from the mask and passed to the transformer blocks, not only the max sequence length across all batch samples for RoPE

dxqb · 2025-11-29T06:56:41Z

src/diffusers/models/controlnets/controlnet_qwenimage.py

+                raise ValueError(f"`txt_seq_lens` must have length {batch_size}, but got {len(txt_seq_lens)} instead.")
+            text_seq_len = max(text_seq_len, max(txt_seq_lens))
+        elif encoder_hidden_states_mask is not None:
+            text_seq_len = max(text_seq_len, int(encoder_hidden_states_mask.sum(dim=1).max().item()))


This only works if the attention mask is in the form of [True, True, True, ..., False, False, False]. While this is the case in the most common use case of text attention masks, it doesn't have to be the case.

If the mask is [True, False, True, False, True, False], self.pos_embed receives an incorrect sequence length

Fix QwenImage txt_seq_lens handling

b547fcf

kashif requested a review from sayakpaul November 23, 2025 18:03

kashif added 2 commits November 23, 2025 18:24

formatting

72a80c6

formatting

88cee8b

sayakpaul requested a review from yiyixuxu November 24, 2025 01:57

dxqb reviewed Nov 29, 2025

View reviewed changes

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Fix QwenImage txt_seq_lens handling #12702

Fix QwenImage txt_seq_lens handling #12702

kashif commented Nov 23, 2025 •

edited

Loading

Uh oh!

HuggingFaceDocBuilderDev commented Nov 23, 2025

Uh oh!

dxqb commented Nov 29, 2025 •

edited

Loading

Uh oh!

dxqb Nov 29, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

Fix QwenImage txt_seq_lens handling #12702

Are you sure you want to change the base?

Fix QwenImage txt_seq_lens handling #12702

Conversation

kashif commented Nov 23, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

What does this PR do?

Before submitting

Who can review?

Uh oh!

HuggingFaceDocBuilderDev commented Nov 23, 2025

Uh oh!

dxqb commented Nov 29, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

dxqb Nov 29, 2025

Choose a reason for hiding this comment

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

kashif commented Nov 23, 2025 •

edited

Loading

dxqb commented Nov 29, 2025 •

edited

Loading