Salvato in:
| Autori principali: | Yu, Wongyun, Seo, Ahyun, Cho, Minsu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2508.10740 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Leveraging 3D Geometric Priors in 2D Rotation Symmetry Detection
di: Seo, Ahyun, et al.
Pubblicazione: (2025)
di: Seo, Ahyun, et al.
Pubblicazione: (2025)
Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
di: Kim, Dongkeun, et al.
Pubblicazione: (2025)
di: Kim, Dongkeun, et al.
Pubblicazione: (2025)
Towards More Practical Group Activity Detection: A New Benchmark and Model
di: Kim, Dongkeun, et al.
Pubblicazione: (2023)
di: Kim, Dongkeun, et al.
Pubblicazione: (2023)
Current Symmetry Group Equivariant Convolution Frameworks for Representation Learning
di: Basheer, Ramzan, et al.
Pubblicazione: (2024)
di: Basheer, Ramzan, et al.
Pubblicazione: (2024)
3D Equivariant Pose Regression via Direct Wigner-D Harmonics Prediction
di: Lee, Jongmin, et al.
Pubblicazione: (2024)
di: Lee, Jongmin, et al.
Pubblicazione: (2024)
Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping
di: Lee, Junmyeong, et al.
Pubblicazione: (2026)
di: Lee, Junmyeong, et al.
Pubblicazione: (2026)
Learning Correlation Structures for Vision Transformers
di: Kim, Manjin, et al.
Pubblicazione: (2024)
di: Kim, Manjin, et al.
Pubblicazione: (2024)
Locality-Aware Zero-Shot Human-Object Interaction Detection
di: Kim, Sanghyun, et al.
Pubblicazione: (2025)
di: Kim, Sanghyun, et al.
Pubblicazione: (2025)
Similarity-Aware Selective State-Space Modeling for Semantic Correspondence
di: Kim, Seungwook, et al.
Pubblicazione: (2025)
di: Kim, Seungwook, et al.
Pubblicazione: (2025)
In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation
di: Kang, Dahyun, et al.
Pubblicazione: (2024)
di: Kang, Dahyun, et al.
Pubblicazione: (2024)
RapidMV: Leveraging Spatio-Angular Representations for Efficient and Consistent Text-to-Multi-View Synthesis
di: Kim, Seungwook, et al.
Pubblicazione: (2025)
di: Kim, Seungwook, et al.
Pubblicazione: (2025)
PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
di: Jung, Seongmin, et al.
Pubblicazione: (2025)
di: Jung, Seongmin, et al.
Pubblicazione: (2025)
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
di: Kim, Seungwook, et al.
Pubblicazione: (2026)
di: Kim, Seungwook, et al.
Pubblicazione: (2026)
Equivariant Latent Alignment via Flow Matching under Group Symmetries
di: Kim, Sunghyun, et al.
Pubblicazione: (2026)
di: Kim, Sunghyun, et al.
Pubblicazione: (2026)
Harnessing the Power of Training-Free Techniques in Text-to-2D Generation for Text-to-3D Generation via Score Distillation Sampling
di: Lee, Junhong, et al.
Pubblicazione: (2025)
di: Lee, Junhong, et al.
Pubblicazione: (2025)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
di: Kim, Manjin, et al.
Pubblicazione: (2026)
di: Kim, Manjin, et al.
Pubblicazione: (2026)
Contrastive Mean-Shift Learning for Generalized Category Discovery
di: Choi, Sua, et al.
Pubblicazione: (2024)
di: Choi, Sua, et al.
Pubblicazione: (2024)
Affostruction: 3D Affordance Grounding with Generative Reconstruction
di: Park, Chunghyun, et al.
Pubblicazione: (2026)
di: Park, Chunghyun, et al.
Pubblicazione: (2026)
Generic Event Boundary Detection via Denoising Diffusion
di: Hwang, Jaejun, et al.
Pubblicazione: (2025)
di: Hwang, Jaejun, et al.
Pubblicazione: (2025)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
di: Kim, Dongwon, et al.
Pubblicazione: (2026)
di: Kim, Dongwon, et al.
Pubblicazione: (2026)
Video Summarization with Large Language Models
di: Lee, Min Jung, et al.
Pubblicazione: (2025)
di: Lee, Min Jung, et al.
Pubblicazione: (2025)
FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
di: Kim, Seungwook, et al.
Pubblicazione: (2025)
di: Kim, Seungwook, et al.
Pubblicazione: (2025)
DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning
di: Lee, Junha, et al.
Pubblicazione: (2026)
di: Lee, Junha, et al.
Pubblicazione: (2026)
VIRD: View-Invariant Representation through Dual-Axis Transformation for Cross-View Pose Estimation
di: Park, Juhye, et al.
Pubblicazione: (2026)
di: Park, Juhye, et al.
Pubblicazione: (2026)
RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos
di: Lee, Junmyeong, et al.
Pubblicazione: (2024)
di: Lee, Junmyeong, et al.
Pubblicazione: (2024)
Learning SO(3)-Invariant Semantic Correspondence via Local Shape Transform
di: Park, Chunghyun, et al.
Pubblicazione: (2024)
di: Park, Chunghyun, et al.
Pubblicazione: (2024)
Online Temporal Action Localization with Memory-Augmented Transformer
di: Song, Youngkil, et al.
Pubblicazione: (2024)
di: Song, Youngkil, et al.
Pubblicazione: (2024)
Edge-Aware Image Manipulation via Diffusion Models with a Novel Structure-Preservation Loss
di: Gong, Minsu, et al.
Pubblicazione: (2026)
di: Gong, Minsu, et al.
Pubblicazione: (2026)
Exploring High-Order Self-Similarity for Video Understanding
di: Kim, Manjin, et al.
Pubblicazione: (2026)
di: Kim, Manjin, et al.
Pubblicazione: (2026)
Few-Shot Pattern Detection via Template Matching and Regression
di: Jo, Eunchan, et al.
Pubblicazione: (2025)
di: Jo, Eunchan, et al.
Pubblicazione: (2025)
Hierarchically Structured Neural Bones for Reconstructing Animatable Objects from Casual Videos
di: Jeon, Subin, et al.
Pubblicazione: (2024)
di: Jeon, Subin, et al.
Pubblicazione: (2024)
ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation
di: Gong, Dayoung, et al.
Pubblicazione: (2024)
di: Gong, Dayoung, et al.
Pubblicazione: (2024)
Naïve Exposure of Generative AI Capabilities Undermines Deepfake Detection
di: Kim, Sunpill, et al.
Pubblicazione: (2026)
di: Kim, Sunpill, et al.
Pubblicazione: (2026)
Classification Matters: Improving Video Action Detection with Class-Specific Attention
di: Lee, Jinsung, et al.
Pubblicazione: (2024)
di: Lee, Jinsung, et al.
Pubblicazione: (2024)
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers
di: Bae, Jongseong, et al.
Pubblicazione: (2024)
di: Bae, Jongseong, et al.
Pubblicazione: (2024)
Projecting Points to Axes: Oriented Object Detection via Point-Axis Representation
di: Zhao, Zeyang, et al.
Pubblicazione: (2024)
di: Zhao, Zeyang, et al.
Pubblicazione: (2024)
MV-SAM: Multi-view Promptable Segmentation using Pointmap Guidance
di: Jeong, Yoonwoo, et al.
Pubblicazione: (2026)
di: Jeong, Yoonwoo, et al.
Pubblicazione: (2026)
Vanilla Group Equivariant Vision Transformer: Simple and Effective
di: Fu, Jiahong, et al.
Pubblicazione: (2026)
di: Fu, Jiahong, et al.
Pubblicazione: (2026)
Affogato: Learning Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
di: Lee, Junha, et al.
Pubblicazione: (2025)
di: Lee, Junha, et al.
Pubblicazione: (2025)
ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors
di: Kim, Minsu, et al.
Pubblicazione: (2025)
di: Kim, Minsu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Leveraging 3D Geometric Priors in 2D Rotation Symmetry Detection
di: Seo, Ahyun, et al.
Pubblicazione: (2025) -
Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
di: Kim, Dongkeun, et al.
Pubblicazione: (2025) -
Towards More Practical Group Activity Detection: A New Benchmark and Model
di: Kim, Dongkeun, et al.
Pubblicazione: (2023) -
Current Symmetry Group Equivariant Convolution Frameworks for Representation Learning
di: Basheer, Ramzan, et al.
Pubblicazione: (2024) -
3D Equivariant Pose Regression via Direct Wigner-D Harmonics Prediction
di: Lee, Jongmin, et al.
Pubblicazione: (2024)