Skip to main content
The TextEncodeHunyuanVideo_ImageToVideo node creates conditioning data for image-to-video generation by combining a text prompt with visual information from a reference image. It uses a CLIP model to process both the text and the image embeddings from a CLIP vision output, then generates tokens that blend these two sources according to the image_interleave setting.

Inputs

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 016b87ead6f7a6ca61eff220e57f59252018cc78e80ec8cff5b83223b8f90f73