XNet v2: Fewer Limitations, Better Results and Greater Universality
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Yanfeng, Li, Lingrui, Wang, Zichen, Liu, Guole, Liu, Ziwen, Yang, Ge |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Partial Channel Network: Compute Fewer, Perform Better
di: Huang, Haiduo, et al.
Pubblicazione: (2025)
di: Huang, Haiduo, et al.
Pubblicazione: (2025)
Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
di: Chen, Shimin, et al.
Pubblicazione: (2024)
di: Chen, Shimin, et al.
Pubblicazione: (2024)
Underwater Camouflaged Object Tracking Meets Vision-Language SAM2
di: Zhang, Chunhui, et al.
Pubblicazione: (2024)
di: Zhang, Chunhui, et al.
Pubblicazione: (2024)
Scaling Artificial Intelligence for Multi-Tumor Early Detection with More Reports, Fewer Masks
di: Bassi, Pedro R. A. S., et al.
Pubblicazione: (2025)
di: Bassi, Pedro R. A. S., et al.
Pubblicazione: (2025)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
All in One: Exploring Unified Vision-Language Tracking with Multi-Modal Alignment
di: Zhang, Chunhui, et al.
Pubblicazione: (2023)
di: Zhang, Chunhui, et al.
Pubblicazione: (2023)
MoiréXNet: Adaptive Multi-Scale Demoiréing with Linear Attention Test-Time Training and Truncated Flow Matching Prior
di: Li, Liangyan, et al.
Pubblicazione: (2025)
di: Li, Liangyan, et al.
Pubblicazione: (2025)
Few and Fewer: Learning Better from Few Examples Using Fewer Base Classes
di: Lafargue, Raphael, et al.
Pubblicazione: (2024)
di: Lafargue, Raphael, et al.
Pubblicazione: (2024)
Awesome Multi-modal Object Tracking
di: Zhang, Chunhui, et al.
Pubblicazione: (2024)
di: Zhang, Chunhui, et al.
Pubblicazione: (2024)
COST: Contrastive One-Stage Transformer for Vision-Language Small Object Tracking
di: Zhang, Chunhui, et al.
Pubblicazione: (2025)
di: Zhang, Chunhui, et al.
Pubblicazione: (2025)
UniHM: Universal Human Motion Generation with Object Interactions in Indoor Scenes
di: Geng, Zichen, et al.
Pubblicazione: (2025)
di: Geng, Zichen, et al.
Pubblicazione: (2025)
From Fewer Samples to Fewer Bits: Reframing Dataset Distillation as Joint Optimization of Precision and Compactness
di: Dinh, My H., et al.
Pubblicazione: (2026)
di: Dinh, My H., et al.
Pubblicazione: (2026)
WebUOT-1M: Advancing Deep Underwater Object Tracking with A Million-Scale Benchmark
di: Zhang, Chunhui, et al.
Pubblicazione: (2024)
di: Zhang, Chunhui, et al.
Pubblicazione: (2024)
PKI: Prior Knowledge-Infused Neural Network for Few-Shot Class-Incremental Learning
di: Baoa, Kexin, et al.
Pubblicazione: (2026)
di: Baoa, Kexin, et al.
Pubblicazione: (2026)
SwiftBrush v2: Make Your One-step Diffusion Model Better Than Its Teacher
di: Dao, Trung, et al.
Pubblicazione: (2024)
di: Dao, Trung, et al.
Pubblicazione: (2024)
Selective Visual Prompting in Vision Mamba
di: Yao, Yifeng, et al.
Pubblicazione: (2024)
di: Yao, Yifeng, et al.
Pubblicazione: (2024)
MRN: Harnessing 2D Vision Foundation Models for Diagnosing Parkinson's Disease with Limited 3D MR Data
di: Shaodong, Ding, et al.
Pubblicazione: (2025)
di: Shaodong, Ding, et al.
Pubblicazione: (2025)
Representing Topological Self-Similarity Using Fractal Feature Maps for Accurate Segmentation of Tubular Structures
di: Huang, Jiaxing, et al.
Pubblicazione: (2024)
di: Huang, Jiaxing, et al.
Pubblicazione: (2024)
Mamba-3D as Masked Autoencoders for Accurate and Data-Efficient Analysis of Medical Ultrasound Videos
di: Zhou, Jiaheng, et al.
Pubblicazione: (2025)
di: Zhou, Jiaheng, et al.
Pubblicazione: (2025)
Efficient-LVSM: Faster, Cheaper, and Better Large View Synthesis Model via Decoupled Co-Refinement Attention
di: Jia, Xiaosong, et al.
Pubblicazione: (2026)
di: Jia, Xiaosong, et al.
Pubblicazione: (2026)
PolarBEVDet: Exploring Polar Representation for Multi-View 3D Object Detection in Bird's-Eye-View
di: Yu, Zichen, et al.
Pubblicazione: (2024)
di: Yu, Zichen, et al.
Pubblicazione: (2024)
TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens
di: Yu, Ya-Qi, et al.
Pubblicazione: (2024)
di: Yu, Ya-Qi, et al.
Pubblicazione: (2024)
Point2RBox-v3: Self-Bootstrapping from Point Annotations via Integrated Pseudo-Label Refinement and Utilization
di: Zhang, Teng, et al.
Pubblicazione: (2025)
di: Zhang, Teng, et al.
Pubblicazione: (2025)
Better Eyes, Better Thoughts: Why Vision Chain-of-Thought Fails in Medicine
di: Wu, Yuan, et al.
Pubblicazione: (2026)
di: Wu, Yuan, et al.
Pubblicazione: (2026)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
di: Li, Kevin Y., et al.
Pubblicazione: (2024)
di: Li, Kevin Y., et al.
Pubblicazione: (2024)
MatchSeg: Towards Better Segmentation via Reference Image Matching
di: Huo, Jiayu, et al.
Pubblicazione: (2024)
di: Huo, Jiayu, et al.
Pubblicazione: (2024)
Does Bigger Mean Better? Comparitive Analysis of CNNs and Biomedical Vision Language Modles in Medical Diagnosis
di: Tong, Ran, et al.
Pubblicazione: (2025)
di: Tong, Ran, et al.
Pubblicazione: (2025)
DiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models
di: Dong, Zhe, et al.
Pubblicazione: (2025)
di: Dong, Zhe, et al.
Pubblicazione: (2025)
Dolphin v1.0 Technical Report
di: Weng, Taohan, et al.
Pubblicazione: (2025)
di: Weng, Taohan, et al.
Pubblicazione: (2025)
NeuFlow v2: Push High-Efficiency Optical Flow To the Limit
di: Zhang, Zhiyong, et al.
Pubblicazione: (2024)
di: Zhang, Zhiyong, et al.
Pubblicazione: (2024)
Learning to Transform Dynamically for Better Adversarial Transferability
di: Zhu, Rongyi, et al.
Pubblicazione: (2024)
di: Zhu, Rongyi, et al.
Pubblicazione: (2024)
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
di: Fu, Ling, et al.
Pubblicazione: (2024)
di: Fu, Ling, et al.
Pubblicazione: (2024)
Deep Neighbor Layer Aggregation for Lightweight Self-Supervised Monocular Depth Estimation
di: Boya, Wang, et al.
Pubblicazione: (2023)
di: Boya, Wang, et al.
Pubblicazione: (2023)
Disentangled Hierarchical VAE for 3D Human-Human Interaction Generation
di: Geng, Zichen, et al.
Pubblicazione: (2026)
di: Geng, Zichen, et al.
Pubblicazione: (2026)
T2I-VeRW: Part-level Fine-grained Perception for Text-to-Image Vehicle Retrieval
di: Wang, Xiao, et al.
Pubblicazione: (2026)
di: Wang, Xiao, et al.
Pubblicazione: (2026)
Compression for Better: A General and Stable Lossless Compression Framework
di: Zhang, Boyang, et al.
Pubblicazione: (2024)
di: Zhang, Boyang, et al.
Pubblicazione: (2024)
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning
di: Ou, Siqu, et al.
Pubblicazione: (2025)
di: Ou, Siqu, et al.
Pubblicazione: (2025)
EvoStreaming: Your Offline Video Model Is a Natively Streaming Assistant
di: Wen, Zichen, et al.
Pubblicazione: (2026)
di: Wen, Zichen, et al.
Pubblicazione: (2026)
Fast Adversarial Training with Weak-to-Strong Spatial-Temporal Consistency in the Frequency Domain on Videos
di: Wang, Songping, et al.
Pubblicazione: (2025)
di: Wang, Songping, et al.
Pubblicazione: (2025)
Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among Instances
di: Yu, Yi, et al.
Pubblicazione: (2025)
di: Yu, Yi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Partial Channel Network: Compute Fewer, Perform Better
di: Huang, Haiduo, et al.
Pubblicazione: (2025) -
Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
di: Chen, Shimin, et al.
Pubblicazione: (2024) -
Underwater Camouflaged Object Tracking Meets Vision-Language SAM2
di: Zhang, Chunhui, et al.
Pubblicazione: (2024) -
Scaling Artificial Intelligence for Multi-Tumor Early Detection with More Reports, Fewer Masks
di: Bassi, Pedro R. A. S., et al.
Pubblicazione: (2025) -
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
di: Ge, Yuyao, et al.
Pubblicazione: (2025)