Gecko: An Efficient Neural Architecture Inherently Processing Sequences with Arbitrary Lengths
Fuente:
arXiv
Guardado en:
| Autores principales: | Ma, Xuezhe, Wen, Shicheng, Jin, Linghao, Acun, Bilge, Lai, Ruihang, Hou, Bohan, Lin, Will, Zhang, Hao, Yang, Songlin, Lee, Ryan, Wu, Mengxi, May, Jonathan, Zettlemoyer, Luke, Wu, Carole-Jean |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond Efficiency: Scaling AI Sustainably
por: Wu, Carole-Jean, et al.
Publicado: (2024)
por: Wu, Carole-Jean, et al.
Publicado: (2024)
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
por: Ma, Xuezhe, et al.
Publicado: (2024)
por: Ma, Xuezhe, et al.
Publicado: (2024)
Unlocking the Potential of Renewable Energy Through Curtailment Prediction
por: Acun, Bilge, et al.
Publicado: (2024)
por: Acun, Bilge, et al.
Publicado: (2024)
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
por: Pepe, Alberto, et al.
Publicado: (2026)
por: Pepe, Alberto, et al.
Publicado: (2026)
Towards Chapter-to-Chapter Context-Aware Literary Translation via Large Language Models
por: Jin, Linghao, et al.
Publicado: (2024)
por: Jin, Linghao, et al.
Publicado: (2024)
Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
por: Bae, Sangmin, et al.
Publicado: (2025)
por: Bae, Sangmin, et al.
Publicado: (2025)
Composer: A Search Framework for Hybrid Neural Architecture Design
por: Acun, Bilge, et al.
Publicado: (2025)
por: Acun, Bilge, et al.
Publicado: (2025)
Light-weight Fine-tuning Method for Defending Adversarial Noise in Pre-trained Medical Vision-Language Models
por: Han, Xu, et al.
Publicado: (2024)
por: Han, Xu, et al.
Publicado: (2024)
CHAI: Clustered Head Attention for Efficient LLM Inference
por: Agarwal, Saurabh, et al.
Publicado: (2024)
por: Agarwal, Saurabh, et al.
Publicado: (2024)
Scalable LLM Reasoning Acceleration with Low-rank Distillation
por: Dong, Harry, et al.
Publicado: (2025)
por: Dong, Harry, et al.
Publicado: (2025)
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
por: Hsia, Samuel, et al.
Publicado: (2023)
por: Hsia, Samuel, et al.
Publicado: (2023)
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
por: Wang, Irene, et al.
Publicado: (2025)
por: Wang, Irene, et al.
Publicado: (2025)
Dived-Jin/Gecko_Sexchromosome: Gecko_sexchromsome
por: jin
Publicado: (2026)
por: jin
Publicado: (2026)
MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models
por: Chiu, Ming-Chang, et al.
Publicado: (2024)
por: Chiu, Ming-Chang, et al.
Publicado: (2024)
Enhancing Gluten‐Free Cake Quality With Germinated and Ungerminated Mung Bean Flours: A Comparative Analysis
por: Sultan Acun
Publicado: (2026)
por: Sultan Acun
Publicado: (2026)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
por: Yang, Songlin, et al.
Publicado: (2024)
por: Yang, Songlin, et al.
Publicado: (2024)
Golay Complementary Sequences of Arbitrary Length and Asymptotic Existence of Hadamard Matrices
por: Du, Cheng, et al.
Publicado: (2024)
por: Du, Cheng, et al.
Publicado: (2024)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
por: Shi, Chufan, et al.
Publicado: (2026)
por: Shi, Chufan, et al.
Publicado: (2026)
Hypoxic adaptation mechanism of polysaccharide from Agaricus bitorquis (Quél.) Sacc.Chaidam on gut microbiota in Tibetan Plateau population based on in vitro model.
por: Qiu, Songlin, et al.
Publicado: (2026)
por: Qiu, Songlin, et al.
Publicado: (2026)
Björck Sequences: Extension to Arbitrary Lengths, Correlation Analysis, and Applications to Wireless Systems
por: Dureppagari, Harish K., et al.
Publicado: (2025)
por: Dureppagari, Harish K., et al.
Publicado: (2025)
PatentEdits: Framing Patent Novelty as Textual Entailment
por: Lee, Ryan, et al.
Publicado: (2024)
por: Lee, Ryan, et al.
Publicado: (2024)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
por: Elhoushi, Mostafa, et al.
Publicado: (2024)
por: Elhoushi, Mostafa, et al.
Publicado: (2024)
Wavefield Correlation Imaging in Arbitrary Media with Inherent Aberration Correction
por: Schoen Jr, Scott, et al.
Publicado: (2025)
por: Schoen Jr, Scott, et al.
Publicado: (2025)
Flora: Effortless Context Construction to Arbitrary Length and Scale
por: Chen, Tianxiang, et al.
Publicado: (2025)
por: Chen, Tianxiang, et al.
Publicado: (2025)
LoopExpose: An Unsupervised Framework for Arbitrary-Length Exposure Correction
por: Li, Ao, et al.
Publicado: (2025)
por: Li, Ao, et al.
Publicado: (2025)
Is Flash Attention Stable?
por: Golden, Alicia, et al.
Publicado: (2024)
por: Golden, Alicia, et al.
Publicado: (2024)
Generative AI Beyond LLMs: System Implications of Multi-Modal Generation
por: Golden, Alicia, et al.
Publicado: (2023)
por: Golden, Alicia, et al.
Publicado: (2023)
Inherently Disordered Auxetic Metamaterials
por: Matteo Montanari, et al.
Publicado: (2026)
por: Matteo Montanari, et al.
Publicado: (2026)
EMO: Frustratingly Easy Progressive Training of Extendable MoE
por: Jin, Linghao, et al.
Publicado: (2026)
por: Jin, Linghao, et al.
Publicado: (2026)
Parallelizing MCMC Across the Sequence Length
por: Zoltowski, David M., et al.
Publicado: (2025)
por: Zoltowski, David M., et al.
Publicado: (2025)
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
por: Kilian, Maciej, et al.
Publicado: (2024)
por: Kilian, Maciej, et al.
Publicado: (2024)
Comparing Hallucination Detection Metrics for Multilingual Generation
por: Kang, Haoqiang, et al.
Publicado: (2024)
por: Kang, Haoqiang, et al.
Publicado: (2024)
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
por: Yasunaga, Michihiro, et al.
Publicado: (2025)
por: Yasunaga, Michihiro, et al.
Publicado: (2025)
(Mis)Fitting: A Survey of Scaling Laws
por: Li, Margaret, et al.
Publicado: (2025)
por: Li, Margaret, et al.
Publicado: (2025)
Transformers are Inherently Succinct
por: Bergsträßer, Pascal, et al.
Publicado: (2025)
por: Bergsträßer, Pascal, et al.
Publicado: (2025)
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
por: Zhou, Chunting, et al.
Publicado: (2024)
por: Zhou, Chunting, et al.
Publicado: (2024)
Functional and Statistical Assessment of Cakes Enriched With Pomegranate Peel and Flower Powders
por: Sultan Acun, et al.
Publicado: (2026)
por: Sultan Acun, et al.
Publicado: (2026)
Optimization of Beyond Diagonal RIS: A Universal Framework Applicable to Arbitrary Architectures
por: Wu, Zheyu, et al.
Publicado: (2024)
por: Wu, Zheyu, et al.
Publicado: (2024)
Giant-Tailed Gecko Optimization Algorithm
por: Zhang, Jincheng
Publicado: (2026)
por: Zhang, Jincheng
Publicado: (2026)
Multiscale Mineralization in the Leopard Gecko Eggshell
por: Joseph Deering, et al.
Publicado: (2024)
por: Joseph Deering, et al.
Publicado: (2024)
Ejemplares similares
-
Beyond Efficiency: Scaling AI Sustainably
por: Wu, Carole-Jean, et al.
Publicado: (2024) -
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
por: Ma, Xuezhe, et al.
Publicado: (2024) -
Unlocking the Potential of Renewable Energy Through Curtailment Prediction
por: Acun, Bilge, et al.
Publicado: (2024) -
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
por: Pepe, Alberto, et al.
Publicado: (2026) -
Towards Chapter-to-Chapter Context-Aware Literary Translation via Large Language Models
por: Jin, Linghao, et al.
Publicado: (2024)