AMC faculty advances agentic AI and video generation research at SIGGRAPH 2026

2026-07-17
Figure showcasing the varied capabilities of UniVidX, a unified multimodal framework proposed by AMC faculty member Prof. Anyi RAO and his team for versatile video generation.

The Division of Arts and Machine Creativity (AMC) at the Hong Kong University of Science and Technology (HKUST) is thrilled to announce the versatile contribution of our faculty members at the SIGGRAPH 2026, one of the world's premier conference and exhibition on computer graphics and interactive techniques, scheduled to be held in Los Angeles, United States from 19–23 July, 2026.

Building on AMC's established strength in AI and generative model researches, as well as Acting Head Prof. Hongbo FU's recent induction as the first Hong Kong scholar to the ACM SIGGRAPH Academy, AMC faculty members shine further at this year's SIGGRAPH conference in the area of agentic AI and video diffusion models (VDM), with three Technical Paper accepted, a major course on multi-agent-driven video synthesis and three technical workshop offered together with distinguished academic collaborators. 

The Division warmly congratulates Prof. Hongbo FU, Prof. Anyi RAO and their collaborators on these remarkable achievements. We look forward to their continued contributions to shaping the future of generative AI, visual computing, and creative technologies on the global stage.

 

Accepted Technical Paper
  1. Controllable Texture Tiling with Transformed RoPE-Enhanced Diffusion Models
    Junrong Huang, Zhiyuan Zhang, Rui Tang, Hongbo Fu, Jing Liao
    Presents a diffusion transformer-based framework for controllable, high-fidelity texture tiling in images, enabling users to precisely adjust texture repetition, scale, and orientation while preserving the scene’s original lighting and geometry. Key innovations include a Coordinate-Transformed Rotary Embedding that controls texture placement through transformed positional embeddings (without degrading the reference texture) and a Disjoint Attention Mask that prevents semantic leakage, resulting in more accurate and faithful texture transfer than existing methods.
  2. CameraSquad: Achieving Content Consistency in Parallel Multi-Trajectory Camera-Controlled Video Generation
    Zhufeng Xu, Xuan Gao, Bailin Deng, Yikang Ding, Xiaoqiang Liu, Haoxian Zhang, Pengfei Wan, Hongbo Fu, Lin Gao
    Presents a new framework for camera-controlled video generation that supports both single-camera trajectories and multiple trajectories generated in parallel, avoiding the viewpoint inconsistencies caused by running separate diffusion-based generations. It employs a dual-mode cross-view attention mechanism that maintains content consistency across different camera views while preserving precise camera control, with experiment-proven strong performance in camera accuracy, cross-view consistency, and overall video quality.
  3. UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
    Houyuan Chen, Hong Li, Xianghao Kong, Tianrui Zhu, Shaocong Xu, Weiqing Xiao, Yuwei Guo, Chongjie Ye, Lvmin Zhang, Hao Zhao, Anyi Rao
    Presents UniVidX, a unified multimodal framework designed to leverage VDM priors to enable versatile video generation, by enabling the model to learn omni-directional conditional generation rather than fixed mappings, preserving the VDM’s strong prior, and sharing keys/values across modalities while maintaining modality-specific queries, to achieve robust generalization capabilities in in-the-wild scenarios, even when trained on limited datasets.

 

Course & Technical Workshop Co-Organization

[Course] AI for Creative Visual Content Generation Editing and Understanding
Organized by Zheng Wei, Yuying Tang, Mia Tang, Anyi Rao
This course introduces a multi-agent framework fusing Generative AI’s creativity with Computer Graphics’ precision. Addressing GenAI’s limitations in data accuracy, it will be demonstrated how agents orchestrate deterministic rendering (e.g., for charts or spatial logic) to guide video synthesis. Attendees will learn to build systems for controllable, precise cinematic and data storytelling.

[Technical Workshop] Human–AI Co-Creation in Generative Art: Graphics Methods, Systems, and Applications
Organized by Ying Jiang, Jiayin Lu, Yin Yang, Hongbo Fu, Chenfanfu Jiang, Yunuo Chen
Recent advances in generative AI are transforming how visual and spatial content is created, enabling new forms of collaboration between humans and intelligent systems. This workshop explores human–AI co-creation in generative art, where AI acts not merely as an automation tool but as a creative partner supporting exploration, iteration, and artistic expression. The workshop brings together researchers, artists, and engineers from computer graphics, machine learning, human–computer interaction, and computational design to examine emerging systems for collaborative creativity. Topics include interactive creative tools, generative models for visual and 3D art, and fabrication-aware design systems that connect digital creation with real-world production.

[Technical Workshop] Asiagraphics Workshop on Neural Graphics Processing
Organized by Young J. Kim, Hongbo Fu, Takeo Igarashi, Ligang Liu, Stefanie Zollmann, Lin Gao
Neural graphics is transforming how we create, represent, and interact with visual content. This workshop brings together leading researchers across geometry, rendering, generative modeling, and embodied AI to present a unified view of this rapidly evolving field. Through a curated sequence of invited talks, the program will guide attendees from geometric foundations and 3D generation to neural rendering, real-time techniques, physics-aware reconstruction, and interaction with the physical world. The primary objective is to explore how neural representations and generative models can be integrated with classical graphics pipelines to enable more scalable, controllable, and physically grounded visual computing.

[Technical Workshop] AI for Creative Visual Content Generation, Editing and Understanding
Organized by Songlin Yang, Gaurav Parmar, Ozgur Kara, Jiaju Ma, Lvmin Zhang, Or Patashnik, Fabian Caba Heilbron, Jun-Yan Zhu, Daniel Cohen-Or, Maneesh Agrawala, Anyi Rao
The 10th installment of the CVEU workshop series marks a transformative shift toward the Next-Generation Creator-Machine Co-Learning Paradigm. Organized by leading academic researchers and the technical architects behind current video generation models, this workshop explores a future where generative AI evolves from a passive tool into a collaborative partner that internalizes both physical laws and artistic intent. The primary objective is to bridge the "semantic gap" between raw visual tokens and the sophisticated language of cinema, focusing on how World Models can balance physical factuality with creative license. One will investigate how machine learning translate abstract user intent into precise model behavior, effectively reconstructing traditional production pipelines for professional filmmakers and artists alike.

 

Figure showcasing the varied capabilities of UniVidX, a unified multimodal framework proposed by AMC faculty member Prof. Anyi RAO and his team for versatile video generation.
G1.png
A display of the capability of the Controllable Texture Tiling framework in rendering textural changes in precise size and orientation steps
G2.png
Figure showing the CameraSquad method with strong consistency across different camera trajectories generated in parallel