StructFormer: Document Structure-based Masked Attention and its Impact on Language Model Pre-Training
Fuente:
arXiv
Guardado en:
| Autores principales: | Ponkshe, Kaustubh, Subramanian, Venkatapathy, Modani, Natwar, Ramakrishnan, Ganesh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GenSco: Can Question Decomposition based Passage Alignment improve Question Answering?
por: Fazili, Barah, et al.
Publicado: (2024)
por: Fazili, Barah, et al.
Publicado: (2024)
GUIDEQ: Framework for Guided Questioning for progressive informational collection and classification
por: Mishra, Priya, et al.
Publicado: (2024)
por: Mishra, Priya, et al.
Publicado: (2024)
FedEx-LoRA: Exact Aggregation for Federated and Efficient Fine-Tuning of Foundation Models
por: Singhal, Raghav, et al.
Publicado: (2024)
por: Singhal, Raghav, et al.
Publicado: (2024)
SPRINT: Script-agnostic Structure Recognition in Tables
por: Kudale, Dhruv, et al.
Publicado: (2025)
por: Kudale, Dhruv, et al.
Publicado: (2025)
A Lightweight Method to Disrupt Memorized Sequences in LLM
por: Prashant, Parjanya Prajakta, et al.
Publicado: (2025)
por: Prashant, Parjanya Prajakta, et al.
Publicado: (2025)
ABBA-Adapters: Efficient and Expressive Fine-Tuning of Foundation Models
por: Singhal, Raghav, et al.
Publicado: (2025)
por: Singhal, Raghav, et al.
Publicado: (2025)
Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study
por: Ponkshe, Kaustubh, et al.
Publicado: (2025)
por: Ponkshe, Kaustubh, et al.
Publicado: (2025)
TEXTRON: Weakly Supervised Multilingual Text Detection through Data Programming
por: Kudale, Dhruv, et al.
Publicado: (2024)
por: Kudale, Dhruv, et al.
Publicado: (2024)
PLATTER: A Page-Level Handwritten Text Recognition System for Indic Scripts
por: Kasuba, Badri Vishal, et al.
Publicado: (2025)
por: Kasuba, Badri Vishal, et al.
Publicado: (2025)
INSITE: labelling medical images using submodular functions and semi-supervised data programming
por: Gautam, Akshat, et al.
Publicado: (2024)
por: Gautam, Akshat, et al.
Publicado: (2024)
Improving Selective Classification with Pairwise Queries for Binary Classification
por: Vardhan, Harsh, et al.
Publicado: (2026)
por: Vardhan, Harsh, et al.
Publicado: (2026)
Struct-X: Enhancing Large Language Models Reasoning with Structured Data
por: Tan, Xiaoyu, et al.
Publicado: (2024)
por: Tan, Xiaoyu, et al.
Publicado: (2024)
Centered Masking for Language-Image Pre-Training
por: Liang, Mingliang, et al.
Publicado: (2024)
por: Liang, Mingliang, et al.
Publicado: (2024)
Masked Structural Growth for 2x Faster Language Model Pre-training
por: Yao, Yiqun, et al.
Publicado: (2023)
por: Yao, Yiqun, et al.
Publicado: (2023)
Fed-SB: A Silver Bullet for Extreme Communication Efficiency and Performance in (Private) Federated LoRA Fine-Tuning
por: Singhal, Raghav, et al.
Publicado: (2025)
por: Singhal, Raghav, et al.
Publicado: (2025)
Analysing The Impact of Sequence Composition on Language Model Pre-Training
por: Zhao, Yu, et al.
Publicado: (2024)
por: Zhao, Yu, et al.
Publicado: (2024)
Initialization using Update Approximation is a Silver Bullet for Extremely Efficient Low-Rank Fine-Tuning
por: Ponkshe, Kaustubh, et al.
Publicado: (2024)
por: Ponkshe, Kaustubh, et al.
Publicado: (2024)
StructLM: Towards Building Generalist Models for Structured Knowledge Grounding
por: Zhuang, Alex, et al.
Publicado: (2024)
por: Zhuang, Alex, et al.
Publicado: (2024)
StructLens: A Structural Lens for Language Models via Maximum Spanning Trees
por: Sakajo, Haruki, et al.
Publicado: (2026)
por: Sakajo, Haruki, et al.
Publicado: (2026)
Fisher Mask Nodes for Language Model Merging
por: K, Thennal D, et al.
Publicado: (2024)
por: K, Thennal D, et al.
Publicado: (2024)
KAUCUS: Knowledge Augmented User Simulators for Training Language Model Assistants
por: Dhole, Kaustubh D.
Publicado: (2024)
por: Dhole, Kaustubh D.
Publicado: (2024)
LEVOS: Leveraging Vocabulary Overlap with Sanskrit to Generate Technical Lexicons in Indian Languages
por: J, Karthika N, et al.
Publicado: (2024)
por: J, Karthika N, et al.
Publicado: (2024)
Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
por: Guo, Xu, et al.
Publicado: (2026)
por: Guo, Xu, et al.
Publicado: (2026)
StructEval: Deepen and Broaden Large Language Model Assessment via Structured Evaluation
por: Cao, Boxi, et al.
Publicado: (2024)
por: Cao, Boxi, et al.
Publicado: (2024)
On the Diversity of Synthetic Data and its Impact on Training Large Language Models
por: Chen, Hao, et al.
Publicado: (2024)
por: Chen, Hao, et al.
Publicado: (2024)
Masked Diffusion Language Models with Frequency-Informed Training
por: Kosmopoulou, Despoina, et al.
Publicado: (2025)
por: Kosmopoulou, Despoina, et al.
Publicado: (2025)
Modular Techniques for Synthetic Long-Context Data Generation in Language Model Training and Evaluation
por: Subramanian, Seganrasan, et al.
Publicado: (2025)
por: Subramanian, Seganrasan, et al.
Publicado: (2025)
Long Context Pre-Training with Lighthouse Attention
por: Peng, Bowen, et al.
Publicado: (2026)
por: Peng, Bowen, et al.
Publicado: (2026)
Legal Documents Drafting with Fine-Tuned Pre-Trained Large Language Model
por: Lin, Chun-Hsien, et al.
Publicado: (2024)
por: Lin, Chun-Hsien, et al.
Publicado: (2024)
StructTest: Benchmarking LLMs' Reasoning through Compositional Structured Outputs
por: Chen, Hailin, et al.
Publicado: (2024)
por: Chen, Hailin, et al.
Publicado: (2024)
StructKV: Preserving the Structural Skeleton for Scalable Long-Context Inference
por: Chen, Zhirui, et al.
Publicado: (2026)
por: Chen, Zhirui, et al.
Publicado: (2026)
HiStruct+: Improving Extractive Text Summarization with Hierarchical Structure Information
por: Ruan, Qian, et al.
Publicado: (2022)
por: Ruan, Qian, et al.
Publicado: (2022)
ProtStructQA: A Denotation Threshold in Protein Structural Reasoning
por: Mandiga, Aravind, et al.
Publicado: (2026)
por: Mandiga, Aravind, et al.
Publicado: (2026)
A Code Comprehension Benchmark for Large Language Models for Code
por: Havare, Jayant, et al.
Publicado: (2025)
por: Havare, Jayant, et al.
Publicado: (2025)
Development of Pre-Trained Transformer-based Models for the Nepali Language
por: Thapa, Prajwal, et al.
Publicado: (2024)
por: Thapa, Prajwal, et al.
Publicado: (2024)
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
por: Maheshwari, Ayush, et al.
Publicado: (2022)
por: Maheshwari, Ayush, et al.
Publicado: (2022)
Power Mechanism: Private Tabular Representation Release for Model Agnostic Consumption
por: Vepakomma, Praneeth, et al.
Publicado: (2025)
por: Vepakomma, Praneeth, et al.
Publicado: (2025)
StructCoh: Structured Contrastive Learning for Context-Aware Text Semantic Matching
por: Xue, Chao, et al.
Publicado: (2025)
por: Xue, Chao, et al.
Publicado: (2025)
MHQA: A Diverse, Knowledge Intensive Mental Health Question Answering Challenge for Language Models
por: Racha, Suraj, et al.
Publicado: (2025)
por: Racha, Suraj, et al.
Publicado: (2025)
Characterizing the Behavior of Training Mamba-based State Space Models on GPUs
por: Baruah, Trinayan, et al.
Publicado: (2025)
por: Baruah, Trinayan, et al.
Publicado: (2025)
Ejemplares similares
-
GenSco: Can Question Decomposition based Passage Alignment improve Question Answering?
por: Fazili, Barah, et al.
Publicado: (2024) -
GUIDEQ: Framework for Guided Questioning for progressive informational collection and classification
por: Mishra, Priya, et al.
Publicado: (2024) -
FedEx-LoRA: Exact Aggregation for Federated and Efficient Fine-Tuning of Foundation Models
por: Singhal, Raghav, et al.
Publicado: (2024) -
SPRINT: Script-agnostic Structure Recognition in Tables
por: Kudale, Dhruv, et al.
Publicado: (2025) -
A Lightweight Method to Disrupt Memorized Sequences in LLM
por: Prashant, Parjanya Prajakta, et al.
Publicado: (2025)