Rethinking the Zigzag Flattening for Image Reading
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Qingsong, Wang, Yi, Zhou, Zhipeng, Miao, Duoqian, Wang, Limin, Qiao, Yu, Zhao, Cairong |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Transformer-Based Person Search with High-Frequency Augmentation and Multi-Wave Mixing
by: Shu, Qilin, et al.
Published: (2025)
by: Shu, Qilin, et al.
Published: (2025)
Perception Activator: An intuitive and portable framework for brain cognitive exploration
by: Xu, Le, et al.
Published: (2025)
by: Xu, Le, et al.
Published: (2025)
Boosting Adversarial Transferability via Commonality-Oriented Gradient Optimization
by: Gao, Yanting, et al.
Published: (2025)
by: Gao, Yanting, et al.
Published: (2025)
Adaptive Discriminative Regularization for Visual Classification
by: Zhao, Qingsong, et al.
Published: (2022)
by: Zhao, Qingsong, et al.
Published: (2022)
RIVER: A Real-Time Interaction Benchmark for Video LLMs
by: Shi, Yansong, et al.
Published: (2026)
by: Shi, Yansong, et al.
Published: (2026)
Similarity Distribution based Membership Inference Attack on Person Re-identification
by: Gao, Junyao, et al.
Published: (2022)
by: Gao, Junyao, et al.
Published: (2022)
ArchMap: Arch-Flattening and Knowledge-Guided Vision Language Model for Tooth Counting and Structured Dental Understanding
by: Zhang, Bohan, et al.
Published: (2025)
by: Zhang, Bohan, et al.
Published: (2025)
DROP: Decouple Re-Identification and Human Parsing with Task-specific Features for Occluded Person Re-identification
by: Dou, Shuguang, et al.
Published: (2024)
by: Dou, Shuguang, et al.
Published: (2024)
Flattening Singular Values of Factorized Convolution for Medical Images
by: Feng, Zexin, et al.
Published: (2024)
by: Feng, Zexin, et al.
Published: (2024)
CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware Diffusion
by: Ji, Chenhao, et al.
Published: (2025)
by: Ji, Chenhao, et al.
Published: (2025)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
by: Xu, Yicheng, et al.
Published: (2025)
by: Xu, Yicheng, et al.
Published: (2025)
Enhancing Text-to-Image Diffusion Transformer via Split-Text Conditioning
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Markovian Scale Prediction: A New Era of Visual Autoregressive Generation
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Asymmetric Masked Distillation for Pre-Training Small Foundation Models
by: Zhao, Zhiyu, et al.
Published: (2023)
by: Zhao, Zhiyu, et al.
Published: (2023)
ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding
by: Wang, Yubin, et al.
Published: (2024)
by: Wang, Yubin, et al.
Published: (2024)
DragNeXt: Rethinking Drag-Based Image Editing
by: Zhou, Yuan, et al.
Published: (2025)
by: Zhou, Yuan, et al.
Published: (2025)
Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts
by: Wang, Yubin, et al.
Published: (2025)
by: Wang, Yubin, et al.
Published: (2025)
FedMinds: Privacy-Preserving Personalized Brain Visual Decoding
by: Bao, Guangyin, et al.
Published: (2024)
by: Bao, Guangyin, et al.
Published: (2024)
Rethinking the Paradigm of Content Constraints in Unpaired Image-to-Image Translation
by: Cai, Xiuding, et al.
Published: (2022)
by: Cai, Xiuding, et al.
Published: (2022)
Flatten Anything: Unsupervised Neural Surface Parameterization
by: Zhang, Qijian, et al.
Published: (2024)
by: Zhang, Qijian, et al.
Published: (2024)
FlattenGPT: Depth Compression for Transformer with Layer Flattening
by: Xu, Ruihan, et al.
Published: (2026)
by: Xu, Ruihan, et al.
Published: (2026)
One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language Models
by: Zhao, Jiale, et al.
Published: (2025)
by: Zhao, Jiale, et al.
Published: (2025)
VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
by: Wang, Zikang, et al.
Published: (2025)
by: Wang, Zikang, et al.
Published: (2025)
DiffPhysBA: Diffusion-based Physical Backdoor Attack against Person Re-Identification in Real-World
by: Sun, Wenli, et al.
Published: (2024)
by: Sun, Wenli, et al.
Published: (2024)
SM$^3$: Self-Supervised Multi-task Modeling with Multi-view 2D Images for Articulated Objects
by: Wang, Haowen, et al.
Published: (2024)
by: Wang, Haowen, et al.
Published: (2024)
VideoMamba: State Space Model for Efficient Video Understanding
by: Li, Kunchang, et al.
Published: (2024)
by: Li, Kunchang, et al.
Published: (2024)
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
by: Li, Kunchang, et al.
Published: (2023)
by: Li, Kunchang, et al.
Published: (2023)
Adaptive Visual Autoregressive Acceleration via Dual-Linkage Entropy Analysis
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
Dynamic and Compressive Adaptation of Transformers From Images to Videos
by: Zhang, Guozhen, et al.
Published: (2024)
by: Zhang, Guozhen, et al.
Published: (2024)
Neural Image Unfolding: Flattening Sparse Anatomical Structures using Neural Fields
by: Rist, Leonhard, et al.
Published: (2024)
by: Rist, Leonhard, et al.
Published: (2024)
Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
by: Li, Yueyan, et al.
Published: (2025)
by: Li, Yueyan, et al.
Published: (2025)
FW-VTON: Flattening-and-Warping for Person-to-Person Virtual Try-on
by: Wang, Zheng, et al.
Published: (2025)
by: Wang, Zheng, et al.
Published: (2025)
Uni$^2$Det: Unified and Universal Framework for Prompt-Guided Multi-dataset 3D Detection
by: Wang, Yubin, et al.
Published: (2024)
by: Wang, Yubin, et al.
Published: (2024)
HPT++: Hierarchically Prompting Vision-Language Models with Multi-Granularity Knowledge Generation and Improved Structure Modeling
by: Wang, Yubin, et al.
Published: (2024)
by: Wang, Yubin, et al.
Published: (2024)
Flatten: Video Action Recognition is an Image Classification task
by: Chen, Junlin, et al.
Published: (2024)
by: Chen, Junlin, et al.
Published: (2024)
LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
by: Zhang, Hongjie, et al.
Published: (2023)
by: Zhang, Hongjie, et al.
Published: (2023)
TRAIL: Transferable Robust Adversarial Images via Latent diffusion
by: Xue, Yuhao, et al.
Published: (2025)
by: Xue, Yuhao, et al.
Published: (2025)
Infrared and Visible Image Fusion with Language-Driven Loss in CLIP Embedding Space
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Faster and Better: Reinforced Collaborative Distillation and Self-Learning for Infrared-Visible Image Fusion
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Rethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation
by: Cheng, Jiaxin, et al.
Published: (2024)
by: Cheng, Jiaxin, et al.
Published: (2024)
Similar Items
-
Transformer-Based Person Search with High-Frequency Augmentation and Multi-Wave Mixing
by: Shu, Qilin, et al.
Published: (2025) -
Perception Activator: An intuitive and portable framework for brain cognitive exploration
by: Xu, Le, et al.
Published: (2025) -
Boosting Adversarial Transferability via Commonality-Oriented Gradient Optimization
by: Gao, Yanting, et al.
Published: (2025) -
Adaptive Discriminative Regularization for Visual Classification
by: Zhao, Qingsong, et al.
Published: (2022) -
RIVER: A Real-Time Interaction Benchmark for Video LLMs
by: Shi, Yansong, et al.
Published: (2026)