FlattenGPT: Depth Compression for Transformer with Layer Flattening
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xu, Ruihan, Guo, Qingpei, Zhu, Yao, Ji, Xiangyang, Yang, Ming, Zhang, Shiliang |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Flatten: Video Action Recognition is an Image Classification task
par: Chen, Junlin, et autres
Publié: (2024)
par: Chen, Junlin, et autres
Publié: (2024)
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
par: Xuan, Shiyu, et autres
Publié: (2023)
par: Xuan, Shiyu, et autres
Publié: (2023)
Flatten Long-Range Loss Landscapes for Cross-Domain Few-Shot Learning
par: Zou, Yixiong, et autres
Publié: (2024)
par: Zou, Yixiong, et autres
Publié: (2024)
PixelGen: Improving Pixel Diffusion with Perceptual Supervision
par: Ma, Zehong, et autres
Publié: (2026)
par: Ma, Zehong, et autres
Publié: (2026)
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment
par: Ma, Ziping, et autres
Publié: (2024)
par: Ma, Ziping, et autres
Publié: (2024)
Rethinking the Zigzag Flattening for Image Reading
par: Zhao, Qingsong, et autres
Publié: (2022)
par: Zhao, Qingsong, et autres
Publié: (2022)
Flatten Anything: Unsupervised Neural Surface Parameterization
par: Zhang, Qijian, et autres
Publié: (2024)
par: Zhang, Qijian, et autres
Publié: (2024)
BiDepth: A Bidirectional-Depth Neural Network for Spatio-Temporal Prediction
par: Ehsani, Sina, et autres
Publié: (2025)
par: Ehsani, Sina, et autres
Publié: (2025)
FW-VTON: Flattening-and-Warping for Person-to-Person Virtual Try-on
par: Wang, Zheng, et autres
Publié: (2025)
par: Wang, Zheng, et autres
Publié: (2025)
3D-IDE: 3D Implicit Depth Emergent
par: Zhang, Chushan, et autres
Publié: (2026)
par: Zhang, Chushan, et autres
Publié: (2026)
M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
par: Guo, Qingpei, et autres
Publié: (2024)
par: Guo, Qingpei, et autres
Publié: (2024)
Flattening Singular Values of Factorized Convolution for Medical Images
par: Feng, Zexin, et autres
Publié: (2024)
par: Feng, Zexin, et autres
Publié: (2024)
Flattening the Parent Bias: Hierarchical Semantic Segmentation in the Poincaré Ball
par: Weber, Simon, et autres
Publié: (2024)
par: Weber, Simon, et autres
Publié: (2024)
FOAA: Flattened Outer Arithmetic Attention For Multimodal Tumor Classification
par: Alwazzan, Omnia, et autres
Publié: (2024)
par: Alwazzan, Omnia, et autres
Publié: (2024)
SGFormer: Spherical Geometry Transformer for 360 Depth Estimation
par: Zhang, Junsong, et autres
Publié: (2024)
par: Zhang, Junsong, et autres
Publié: (2024)
HOTVCOM: Generating Buzzworthy Comments for Videos
par: Chen, Yuyan, et autres
Publié: (2024)
par: Chen, Yuyan, et autres
Publié: (2024)
Flatten The Complex: Joint B-Rep Generation via Compositional $k$-Cell Particles
par: Lu, Junran, et autres
Publié: (2026)
par: Lu, Junran, et autres
Publié: (2026)
DenseFormer: Learning Dense Depth Map from Sparse Depth and Image via Conditional Diffusion Model
par: Yuan, Ming, et autres
Publié: (2025)
par: Yuan, Ming, et autres
Publié: (2025)
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
par: Chen, Sili, et autres
Publié: (2025)
par: Chen, Sili, et autres
Publié: (2025)
DepthDark: Robust Monocular Depth Estimation for Low-Light Environments
par: Zeng, Longjian, et autres
Publié: (2025)
par: Zeng, Longjian, et autres
Publié: (2025)
UPDP: A Unified Progressive Depth Pruner for CNN and Vision Transformer
par: Liu, Ji, et autres
Publié: (2024)
par: Liu, Ji, et autres
Publié: (2024)
Neural Image Unfolding: Flattening Sparse Anatomical Structures using Neural Fields
par: Rist, Leonhard, et autres
Publié: (2024)
par: Rist, Leonhard, et autres
Publié: (2024)
Model Compression using Progressive Channel Pruning
par: Guo, Jinyang, et autres
Publié: (2025)
par: Guo, Jinyang, et autres
Publié: (2025)
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
par: Zhang, Qinglin, et autres
Publié: (2024)
par: Zhang, Qinglin, et autres
Publié: (2024)
Rethinking Cross-Layer Information Routing in Diffusion Transformers
par: Xu, Chao, et autres
Publié: (2026)
par: Xu, Chao, et autres
Publié: (2026)
ArchMap: Arch-Flattening and Knowledge-Guided Vision Language Model for Tooth Counting and Structured Dental Understanding
par: Zhang, Bohan, et autres
Publié: (2025)
par: Zhang, Bohan, et autres
Publié: (2025)
Incomplete Multi-View Multi-Label Classification via Shared Codebook and Fused-Teacher Self-Distillation
par: Yan, Xu, et autres
Publié: (2026)
par: Yan, Xu, et autres
Publié: (2026)
Layer- and Timestep-Adaptive Differentiable Token Compression Ratios for Efficient Diffusion Transformers
par: You, Haoran, et autres
Publié: (2024)
par: You, Haoran, et autres
Publié: (2024)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
par: Zheng, Naishan, et autres
Publié: (2025)
par: Zheng, Naishan, et autres
Publié: (2025)
Global Context Compression with Interleaved Vision-Text Transformation
par: Jiao, Dian, et autres
Publié: (2026)
par: Jiao, Dian, et autres
Publié: (2026)
Multi-Prompt with Depth Partitioned Cross-Modal Learning
par: Tian, Yingjie, et autres
Publié: (2023)
par: Tian, Yingjie, et autres
Publié: (2023)
Advancing Depth Anything Model for Unsupervised Monocular Depth Estimation in Endoscopy
par: Li, Bojian, et autres
Publié: (2024)
par: Li, Bojian, et autres
Publié: (2024)
RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers
par: Xu, Xuwei, et autres
Publié: (2025)
par: Xu, Xuwei, et autres
Publié: (2025)
Compressing Vision Transformers in Geospatial Transfer Learning with Manifold-Constrained Optimization
par: Snyder, Thomas, et autres
Publié: (2026)
par: Snyder, Thomas, et autres
Publié: (2026)
Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization
par: Zhu, Xuanyu, et autres
Publié: (2026)
par: Zhu, Xuanyu, et autres
Publié: (2026)
CLQ: Cross-Layer Guided Orthogonal-based Quantization for Diffusion Transformers
par: Liu, Kai, et autres
Publié: (2025)
par: Liu, Kai, et autres
Publié: (2025)
Deep Neighbor Layer Aggregation for Lightweight Self-Supervised Monocular Depth Estimation
par: Boya, Wang, et autres
Publié: (2023)
par: Boya, Wang, et autres
Publié: (2023)
GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
par: Yang, Jiahao, et autres
Publié: (2026)
par: Yang, Jiahao, et autres
Publié: (2026)
Towards Open-world Generalized Deepfake Detection: General Feature Extraction via Unsupervised Domain Adaptation
par: Guo, Midou, et autres
Publié: (2025)
par: Guo, Midou, et autres
Publié: (2025)
RDFC-GAN: RGB-Depth Fusion CycleGAN for Indoor Depth Completion
par: Wang, Haowen, et autres
Publié: (2023)
par: Wang, Haowen, et autres
Publié: (2023)
Documents similaires
-
Flatten: Video Action Recognition is an Image Classification task
par: Chen, Junlin, et autres
Publié: (2024) -
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
par: Xuan, Shiyu, et autres
Publié: (2023) -
Flatten Long-Range Loss Landscapes for Cross-Domain Few-Shot Learning
par: Zou, Yixiong, et autres
Publié: (2024) -
PixelGen: Improving Pixel Diffusion with Perceptual Supervision
par: Ma, Zehong, et autres
Publié: (2026) -
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment
par: Ma, Ziping, et autres
Publié: (2024)