UniWorld @ ECCV 2026

UniWorld Workshop @ ECCV2026

Universal Representations for Perception, Reasoning, and World Modeling

Important Dates Deadlines 23:59 AoE

  1. July 1, 2026July 8, 2026 Archival Submission Deadline
  2. Non-archival Submission Deadline
  3. July 20, 2026Aug 3, 2026 Archival Notification
  4. Non-archival Notification
  5. Camera-ready
  6. Workshop Date

Overview

Computer vision is shifting from task-specific pipelines to general-purpose multimodal foundation models. Yet current systems remain fragmented: perception models recognize, generative models synthesize, and reasoning often occurs mainly in language space. This separation limits scalability, transferability, and holistic scene understanding.

Recent progress in multimodal large models, neural scene representations, video foundation models, and embodied world models suggests these directions are converging. UniWorld targets this moment by promoting universal representations that unify perception, reasoning, generation, and interaction within a coherent framework.

Despite rapid advances, core challenges remain: designing architectures for heterogeneous data, mitigating task interference, enabling compositional reasoning, and achieving robust generalization across domains. By bringing together researchers across foundation models, multimodal learning, and world modeling, UniWorld aims to catalyze principled approaches toward generalizable visual intelligence.

Call For Papers

Core topics

  1. Scalable visual foundation representations
  2. Vision-language and beyond multimodality
  3. Unified generation + understanding paradigms
  4. Universal 3D/4D scene modeling
  5. World modeling and embodied intelligence
  6. Transfer, continual learning, generalization
  7. Emerging trends and open challenges

Submission Guidelines

  1. Submit your paper via OpenReview. Submissions must follow the ECCV 2026 Submission Policy.
  2. Submission tracks Archival: accepted papers will be included in the ECCV proceedings. Non-archival: accepted papers will not be included in the proceedings, so we welcome submissions that have been accepted by or are under review at other venues.
  3. Prepare submissions using the ECCV 2026 Author Kit for LaTeX. Note: For the supplementary material, please append it directly to the main PDF submission.
  4. Papers submitted to the workshop will be reviewed in a double-blind process. All accepted papers will be presented in a poster session.
  • Paper LengthMaximum 14 pages, excluding references and supplementary material

Schedule

Time Session Details
08:45-09:00WelcomeOpening remarks
09:00-09:40Invited Speaker 1Dima Damen: Universal Representations in Egocentric Vision
09:40-10:20Invited Speaker 2Mike Shou: What’s next for video AI: From Multimodal to Embodied AI
10:20-10:40Coffee Break / Poster Session-
10:40-11:20Invited Speaker 3Elahe Arani: Robust by Design: World Models for Embodied AI
11:20-12:00Invited Speaker 4Angela Dai: Towards Interactable 3D Spaces
12:00-13:00Lunch / Poster Session-
13:00-13:40Invited Speaker 5Andreas Geiger: Opening the Black Box of Generation and Reconstruction
13:40-14:20Invited Speaker 6Chen Change Loy: From Streaming 3D to Queryable 4D
14:20-15:00Invited Speaker 7Kosta Derpanis: In Search of Universal Concepts
15:00-15:20Coffee Break / Poster Session-
15:20-16:00Invited Speaker 8Pascal Mettes: Hyperbolic geometry as the foundation of unified learning
16:00-16:10Oral Presentation 1Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning
16:10-16:20Oral Presentation 2A Further Leap on the Battle Against Dataset Bias to Improve Generative Self-supervised Learning from Multi-source Data
16:20-16:30Oral Presentation 3Unified Video Dense Prediction from Disjoint Data
16:30-16:40Oral Presentation 4Vero: An Open RL Recipe for General Visual Reasoning
16:40-16:50Oral Presentation 5Slots, Transitions, Loops: Learning Composable World Models for ARC
16:50-17:00Oral Presentation 6Large-scale Pre-training for Grounded Video Caption Generation
17:00-17:10Oral Presentation 7EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
17:10-17:20Oral Presentation 8Prospective Geometric Anchoring for Stable and Controllable Long-Horizon Driving World Models
17:20-17:30Awards & ClosingClosing remarks

Invited Speakers

Organizers

Advisory Board

UniWorld @ ECCV 2026

Accepted Papers

32 papers have been accepted. All accepted papers will be presented during the workshop poster session.

  1. PosterThree Necessary Principles for Self-Supervised Visual Representation Learning
  2. OralA Further Leap on the Battle Against Dataset Bias to Improve Generative Self-supervised Learning from Multi-source Data
  3. OralSlots, Transitions, Loops: Learning Composable World Models for ARC
  4. PosterAquaVision3D: Physically-Consistent Medium Modelling for Underwater Gaussian Splatting
  5. PosterBRo-JEPA: Learning Modular Transformations in Latent Space
  6. PosterSelf-Routed Tensor Adapters for Parameter-Efficient Universal Visual Adaptation
  7. PosterFrom Static Perception to Physical Grounding: A Survey of Visual Reasoning in Physical AI
  8. PosterField-Operator World Models: Learning Transferable Scene Dynamics as Operators over Moving Primitives
  9. PosterRepresentation Forcing for Bottleneck-Free Unified Multimodal Models
  10. PosterUnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
  11. PosterWhy Does RLVR Shrink the Reasoning Boundary? A Two-Mode Theory of Pass@$k$ Inversion and a Per-Problem Fix
  12. PosterDVA-CLIP: Zero-Feature Denoising and Vision-Side Attention Adaptation for Anomaly Detection
  13. PosterGrounded-Dreamer: Robot Learning from World Model Synthetic Data with Physical Grounding
  14. PosterNot All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles
  15. PosterOrbis 2: A Hierarchical World Model for Driving
  16. PosterDoCoG: Mask-based Multi-Type Grounded Chain-of-Thought for Document QA
  17. PosterWhat Moves? Context-Aware Localized Latent Actions
  18. PosterSeeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
  19. PosterRethinking the Backbone in Model-Inversion-Based Exemplar-Free Continual Learning
  20. PosterDINOcular: Self-Supervised Visuospatial Representations
  21. PosterPseudo-Hilbert Masking for I-JEPA Pre-training
  22. PosterSAM3-WM: World Models on Segmentation Features enable Mask-Guided Planning
  23. OralPersistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning
  24. PosterReSink: Stop Words to Improve Training-Free Referring Segmentation
  25. OralEgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
  26. OralProspective Geometric Anchoring for Stable and Controllable Long-Horizon Driving World Models
  27. OralVero: An Open RL Recipe for General Visual Reasoning
  28. PosterRevisiting Cross-View Completion: Self-Supervised Pre-Training via Reconstruction Error Comparison
  29. OralLarge-scale Pre-training for Grounded Video Caption Generation
  30. PosterVulnerability to Multimodal Jailbreak and Prompt Injection Attacks under Activation Quantization
  31. PosterRepresentations Before Pixels: Semantics-Guided Hierarchical Video Prediction
  32. OralUnified Video Dense Prediction from Disjoint Data