ActFormer: Scalable Collaborative Perception via Active Queries
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Suozhi, Zhang, Juexiao, Li, Yiming, Feng, Chen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tell Me Where You Are: Multimodal LLMs Meet Place Recognition
by: Lyu, Zonglin, et al.
Published: (2024)
by: Lyu, Zonglin, et al.
Published: (2024)
HeatV2X: Scalable Heterogeneous Collaborative Perception via Efficient Alignment and Interaction
by: Zhao, Yueran, et al.
Published: (2025)
by: Zhao, Yueran, et al.
Published: (2025)
INSTINCT: Instance-Level Interaction Architecture for Query-Based Collaborative Perception
by: Xu, Yunjiang, et al.
Published: (2025)
by: Xu, Yunjiang, et al.
Published: (2025)
Act2See: Emergent Active Visual Perception for Video Reasoning
by: Ma, Martin Q., et al.
Published: (2026)
by: Ma, Martin Q., et al.
Published: (2026)
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
by: Fang, Irving, et al.
Published: (2025)
by: Fang, Irving, et al.
Published: (2025)
Query Nearby: Offset-Adjusted Mask2Former enhances small-organ segmentation
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
CATNet: Collaborative Alignment and Transformation Network for Cooperative Perception
by: Chen, Gong, et al.
Published: (2026)
by: Chen, Gong, et al.
Published: (2026)
Self-Localized Collaborative Perception
by: Ni, Zhenyang, et al.
Published: (2024)
by: Ni, Zhenyang, et al.
Published: (2024)
RayFormer: Improving Query-Based Multi-Camera 3D Object Detection via Ray-Centric Strategies
by: Chu, Xiaomeng, et al.
Published: (2024)
by: Chu, Xiaomeng, et al.
Published: (2024)
Multiview Scene Graph
by: Zhang, Juexiao, et al.
Published: (2024)
by: Zhang, Juexiao, et al.
Published: (2024)
ViewFormer: Exploring Spatiotemporal Modeling for Multi-View 3D Occupancy Perception via View-Guided Transformers
by: Li, Jinke, et al.
Published: (2024)
by: Li, Jinke, et al.
Published: (2024)
ActFER: Agentic Facial Expression Recognition via Active Tool-Augmented Visual Reasoning
by: Liu, Shifeng, et al.
Published: (2026)
by: Liu, Shifeng, et al.
Published: (2026)
SplatFormer: Point Transformer for Robust 3D Gaussian Splatting
by: Chen, Yutong, et al.
Published: (2024)
by: Chen, Yutong, et al.
Published: (2024)
CurveFormer++: 3D Lane Detection by Curve Propagation with Temporal Curve Queries and Attention
by: Bai, Yifeng, et al.
Published: (2024)
by: Bai, Yifeng, et al.
Published: (2024)
Active Visual Perception: Opportunities and Challenges
by: Li, Yian, et al.
Published: (2025)
by: Li, Yian, et al.
Published: (2025)
V2X-PC: Vehicle-to-everything Collaborative Perception via Point Cluster
by: Liu, Si, et al.
Published: (2024)
by: Liu, Si, et al.
Published: (2024)
Collaborative Multi-Object Tracking with Conformal Uncertainty Propagation
by: Su, Sanbao, et al.
Published: (2023)
by: Su, Sanbao, et al.
Published: (2023)
CoDS: Enhancing Collaborative Perception in Heterogeneous Scenarios via Domain Separation
by: Han, Yushan, et al.
Published: (2025)
by: Han, Yushan, et al.
Published: (2025)
STAMP: Scalable Task And Model-agnostic Collaborative Perception
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
SonoSelect: Efficient Ultrasound Perception via Active Probe Exploration
by: Zhang, Yixin, et al.
Published: (2026)
by: Zhang, Yixin, et al.
Published: (2026)
Quantized Prompt for Efficient Generalization of Vision-Language Models
by: Hao, Tianxiang, et al.
Published: (2024)
by: Hao, Tianxiang, et al.
Published: (2024)
WD-FQDet: Multispectral Detection Transformer via Wavelet Decomposition and Frequency-aware Query Learning
by: Yang, Chunjin, et al.
Published: (2026)
by: Yang, Chunjin, et al.
Published: (2026)
V2VLoc: Robust GNSS-Free Collaborative Perception via LiDAR Localization
by: Lin, Wenkai, et al.
Published: (2025)
by: Lin, Wenkai, et al.
Published: (2025)
CoSDH: Communication-Efficient Collaborative Perception via Supply-Demand Awareness and Intermediate-Late Hybridization
by: Xu, Junhao, et al.
Published: (2025)
by: Xu, Junhao, et al.
Published: (2025)
O2Former:Direction-Aware and Multi-Scale Query Enhancement for SAR Ship Instance Segmentation
by: Gao, F., et al.
Published: (2025)
by: Gao, F., et al.
Published: (2025)
Pragmatic Communication in Multi-Agent Collaborative Perception
by: Hu, Yue, et al.
Published: (2024)
by: Hu, Yue, et al.
Published: (2024)
Rate-Distortion Optimized Communication for Collaborative Perception
by: Liu, Genjia, et al.
Published: (2025)
by: Liu, Genjia, et al.
Published: (2025)
TextFormer: A Query-based End-to-End Text Spotter with Mixed Supervision
by: Zhai, Yukun, et al.
Published: (2023)
by: Zhai, Yukun, et al.
Published: (2023)
WhisperNet: A Scalable Solution for Bandwidth-Efficient Collaboration
by: Chen, Gong, et al.
Published: (2026)
by: Chen, Gong, et al.
Published: (2026)
See, Act, Adapt: Active Perception for Unsupervised Cross-Domain Visual Adaptation via Personalized VLM-Guided Agent
by: Tang, Tianci, et al.
Published: (2026)
by: Tang, Tianci, et al.
Published: (2026)
CoopDETR: A Unified Cooperative Perception Framework for 3D Detection via Object Query
by: Wang, Zhe, et al.
Published: (2025)
by: Wang, Zhe, et al.
Published: (2025)
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
by: Wang, Beichen, et al.
Published: (2024)
by: Wang, Beichen, et al.
Published: (2024)
SparseCoop: Cooperative Perception with Kinematic-Grounded Queries
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
by: Gou, Yunhao, et al.
Published: (2025)
by: Gou, Yunhao, et al.
Published: (2025)
ZIP: Scalable Crowd Counting via Zero-Inflated Poisson Modeling
by: Ma, Yiming, et al.
Published: (2025)
by: Ma, Yiming, et al.
Published: (2025)
QORT-Former: Query-optimized Real-time Transformer for Understanding Two Hands Manipulating Objects
by: Ismayilzada, Elkhan, et al.
Published: (2025)
by: Ismayilzada, Elkhan, et al.
Published: (2025)
DATA: Domain-And-Time Alignment for High-Quality Feature Fusion in Collaborative Perception
by: Tian, Chengchang, et al.
Published: (2025)
by: Tian, Chengchang, et al.
Published: (2025)
VistaFormer: Scalable Vision Transformers for Satellite Image Time Series Segmentation
by: MacDonald, Ezra, et al.
Published: (2024)
by: MacDonald, Ezra, et al.
Published: (2024)
Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
by: Wang, Guokang, et al.
Published: (2024)
by: Wang, Guokang, et al.
Published: (2024)
Divide and Conquer: Improving Multi-Camera 3D Perception with 2D Semantic-Depth Priors and Input-Dependent Queries
by: Song, Qi, et al.
Published: (2024)
by: Song, Qi, et al.
Published: (2024)
Similar Items
-
Tell Me Where You Are: Multimodal LLMs Meet Place Recognition
by: Lyu, Zonglin, et al.
Published: (2024) -
HeatV2X: Scalable Heterogeneous Collaborative Perception via Efficient Alignment and Interaction
by: Zhao, Yueran, et al.
Published: (2025) -
INSTINCT: Instance-Level Interaction Architecture for Query-Based Collaborative Perception
by: Xu, Yunjiang, et al.
Published: (2025) -
Act2See: Emergent Active Visual Perception for Video Reasoning
by: Ma, Martin Q., et al.
Published: (2026) -
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
by: Fang, Irving, et al.
Published: (2025)