Hyp2Former: Hierarchy-Aware Hyperbolic Embeddings for Open-Set Panoptic Segmentation

TL;DR

Hyp2Former employs hyperbolic hierarchical embeddings to improve open-set panoptic segmentation, achieving state-of-the-art results.

cs.CV 🔴 Advanced 2026-05-04 41 views
Yao Lu Rohit Mohan Florian Drews Yakov Miron Abhinav Valada
computer vision hierarchical modeling hyperbolic space open-set segmentation deep learning

Key Findings

Methodology

This paper introduces Hyp2Former, which extends Mask2Former by projecting query embeddings into hyperbolic space. It employs a hierarchical proxy-anchor loss that aligns embeddings with class and ancestor proxies, capturing semantic hierarchies. The model optimizes classification and mask prediction in Euclidean space, while encoding semantic relations in hyperbolic space. During inference, instances close to higher-level proxies are identified as unknowns, enabling robust detection without explicit unknown modeling. The approach leverages hyperbolic geometry's exponential volume growth to represent class hierarchies efficiently, improving generalization to unseen categories.

Key Results

  • On MS COCO, Cityscapes, and Lost&Found, Hyp2Former surpasses prior methods, with a PQ of 12.15 on unknown detection, outperforming PoDS and U3HS. It achieves a 8% PQ improvement over previous state-of-the-art, with significant gains in recognition quality (RQ). The model maintains high performance on known classes, with only a minor PQ drop (~5) in open-set settings, demonstrating robustness.
  • In Lost&Found, the model achieves the highest PQ (12.15), indicating superior unknown object detection. On Cityscapes, it retains stable performance with a PQ drop of only 5.29, confirming robustness in urban scenes. On COCO, it effectively detects unknowns while preserving known class accuracy, validating its generalization.
  • Ablation studies show that hierarchical hyperbolic embeddings and proxy-anchor loss contribute significantly to performance. The model's ability to represent semantic hierarchies improves detection of out-of-distribution objects across diverse scenarios.

Significance

This work advances open-world perception by integrating hierarchical semantic structures into dense prediction models. Using hyperbolic space enables compact, low-distortion representations of class hierarchies, addressing limitations of Euclidean embeddings. The approach enhances robustness against distribution shifts and unseen categories, crucial for autonomous systems operating in real-world, unpredictable environments. It bridges the gap between structured semantic understanding and dense scene perception, paving the way for more intelligent, adaptable perception systems in robotics and autonomous driving.

Technical Contribution

The paper introduces a novel hyperbolic hierarchical embedding framework with a hierarchical proxy-anchor loss, enabling continuous semantic relation encoding. It extends Mask2Former by projecting query embeddings into hyperbolic space, capturing class and ancestor relationships. The model maintains classification and mask quality in Euclidean space, while leveraging hyperbolic geometry for semantic structure. This dual-space optimization offers improved generalization to unseen categories, a significant departure from prior flat or pseudo-label-based approaches. The method also demonstrates how hyperbolic geometry's exponential volume growth naturally encodes class hierarchies, providing theoretical and practical benefits.

Novelty

This is the first work to incorporate hyperbolic space for dense open-set panoptic segmentation, explicitly modeling semantic hierarchies at the instance level. Unlike previous methods relying on uncertainty or external data, Hyp2Former learns structured embeddings directly, enabling reliable detection of unknown objects without additional supervision. The hierarchical proxy-anchor loss is a key innovation, promoting semantic abstraction and robustness. This approach fundamentally shifts how semantic relationships are integrated into dense prediction tasks, opening new avenues for open-world scene understanding.

Limitations

  • Computational overhead from hyperbolic operations may hinder real-time deployment, especially in high-resolution scenarios.
  • Dependence on predefined hierarchies limits adaptability to dynamic or evolving class structures, necessitating future work on adaptive hierarchy learning.
  • Performance may degrade in highly cluttered or occluded scenes, where semantic cues are ambiguous, indicating room for integrating contextual reasoning.

Future Work

Future directions include developing adaptive hierarchy learning mechanisms, integrating multi-modal data (e.g., language, LiDAR), and optimizing hyperbolic computations for real-time applications. Exploring self-supervised or unsupervised hierarchy induction could further enhance robustness and scalability, enabling models to operate effectively in diverse, unseen environments.

AI Executive Summary

In the realm of scene understanding, panoptic segmentation has achieved remarkable success in closed-world settings, where all categories are known and fixed during training. However, real-world applications such as autonomous driving and robotics demand models that can recognize and segment unseen objects—an open challenge. Traditional approaches often treat categories as flat, ignoring the rich semantic hierarchies that connect different classes, limiting their ability to generalize to unknown objects.

This paper introduces Hyp2Former, a novel framework that leverages hyperbolic geometry to encode semantic hierarchies directly into the embedding space. By projecting query embeddings into hyperbolic space and employing a hierarchical proxy-anchor loss, the model captures the relationships among classes and their ancestors, enabling it to recognize objects outside the training distribution. The architecture extends Mask2Former, maintaining high-quality mask and classification predictions in Euclidean space while structuring semantic relations hyperbolically.

Experimental results across datasets like MS COCO, Cityscapes, and Lost&Found demonstrate that Hyp2Former outperforms existing methods, achieving a PQ of 12.15 on unknown detection—an 8% improvement over prior state-of-the-art. It maintains robust performance on known classes, with minimal PQ degradation in open-set conditions, confirming its stability and generalization. The approach's core innovation lies in the integration of hyperbolic geometry and hierarchical modeling, addressing the limitations of flat label spaces and enabling more reliable open-world perception.

This work significantly advances the field by providing a theoretically grounded, practically effective method for dense scene understanding in open environments. Its ability to detect and segment unseen objects without additional supervision or external data opens new horizons for autonomous systems, making them safer and more adaptable. Future research will focus on adaptive hierarchies, multi-modal integration, and computational efficiency, further pushing the boundaries of open-set scene understanding.

Deep Dive

Plain Language Accessible to non-experts

想象你在一个大工厂里工作,工厂里有各种不同的机器,比如切割机、焊接机、包装机。每种机器都有自己的类别,但有时候会出现一些新机器,你还没有见过。工厂的管理系统希望能快速识别这些新机器,知道它们大致属于“机械”这个大类别,但又不能确定具体是哪一种。为了做到这一点,系统学习了每个类别之间的关系,比如“焊接机”比“机械”更具体,但又比“工业机器人”更宽泛。它用一种特殊的空间,把这些类别和关系用点和线表示出来,距离越近,代表关系越紧密。这样,即使遇到新机器,系统也能根据它和已知类别的关系,判断它属于“机械”类别,而不是完全陌生。这就像你在学校里认识很多朋友,即使遇到新同学,只要他们和你的朋友关系很近,你就知道他们属于“朋友”这个大组,而不用每次都重新认识一遍。这个方法让计算机变得更聪明,能更好地理解世界,即使面对新事物也能做出合理判断。

ELI14 Explained like you're 14

想象你在学校里,有很多不同的班级,比如数学班、英语班、体育班。每个班级里有很多学生,但有时候会遇到新同学,不知道他们属于哪个班。这个系统就像一个聪明的老师,他不仅知道每个学生的具体班级,还知道这些班级之间的关系,比如“数学班”和“科学班”关系很近,而“体育班”和“音乐班”关系也很近。老师用一种特殊的地图,把所有班级和学生都放在上面,距离越近,代表关系越紧密。这样,当新学生出现时,老师可以根据他们和已知班级的距离,判断他们可能属于哪个大类别,比如“理科学生”或“文科学生”,即使没有见过他们的具体名字。这个方法让老师能更聪明地认识新朋友,不仅知道他们是谁,还知道他们和其他朋友的关系。就像这样,Hyp2Former用数学空间把类别关系画出来,帮助计算机更好地理解世界,特别是遇到新事物时也能做出合理判断。

Abstract

Recognizing unknown objects is crucial for safety-critical applications such as autonomous driving and robotics. Open-Set Panoptic Segmentation (OPS) aims to segment known thing and stuff classes while identifying valid unknown objects as separate instances. Prior OPS approaches largely treat known categories as a flat label set, ignoring the semantic hierarchy that provides valuable structural priors for distinguishing unknown objects from in-distribution classes. In this work, we propose Hyp2Former, an end-to-end framework for OPS that does not require explicit modeling of unknowns during training, and instead learns hierarchical semantic similarities continuously in hyperbolic space. By explicitly encoding hierarchical relationships among known categories, the model learns a structured embedding space that captures multiple levels of semantic abstraction. As a result, unknown objects that cannot be confidently classified as known categories still remain in close proximity to higher-level concepts (e.g., an unknown animal remains closer to "animal" or "object" than to unrelated concepts such as "electronics" or "stuff") and can therefore be reliably detected, even if their fine-grained category was not represented during training. Empirical evaluations across multiple public datasets such as MS COCO, Cityscapes, and Lost&Found demonstrate that Hyp2Former outperforms existing methods on OPS, achieving the best balance between unknown object discovery and in-distribution robustness.

cs.CV cs.AI cs.RO