StructFormer: Document Structure-based Masked Attention and its Impact on Language Model Pre-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Ponkshe, Kaustubh, Subramanian, Venkatapathy, Modani, Natwar, Ramakrishnan, Ganesh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GenSco: Can Question Decomposition based Passage Alignment improve Question Answering?
by: Fazili, Barah, et al.
Published: (2024)
by: Fazili, Barah, et al.
Published: (2024)
GUIDEQ: Framework for Guided Questioning for progressive informational collection and classification
by: Mishra, Priya, et al.
Published: (2024)
by: Mishra, Priya, et al.
Published: (2024)
FedEx-LoRA: Exact Aggregation for Federated and Efficient Fine-Tuning of Foundation Models
by: Singhal, Raghav, et al.
Published: (2024)
by: Singhal, Raghav, et al.
Published: (2024)
SPRINT: Script-agnostic Structure Recognition in Tables
by: Kudale, Dhruv, et al.
Published: (2025)
by: Kudale, Dhruv, et al.
Published: (2025)
A Lightweight Method to Disrupt Memorized Sequences in LLM
by: Prashant, Parjanya Prajakta, et al.
Published: (2025)
by: Prashant, Parjanya Prajakta, et al.
Published: (2025)
ABBA-Adapters: Efficient and Expressive Fine-Tuning of Foundation Models
by: Singhal, Raghav, et al.
Published: (2025)
by: Singhal, Raghav, et al.
Published: (2025)
Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study
by: Ponkshe, Kaustubh, et al.
Published: (2025)
by: Ponkshe, Kaustubh, et al.
Published: (2025)
TEXTRON: Weakly Supervised Multilingual Text Detection through Data Programming
by: Kudale, Dhruv, et al.
Published: (2024)
by: Kudale, Dhruv, et al.
Published: (2024)
PLATTER: A Page-Level Handwritten Text Recognition System for Indic Scripts
by: Kasuba, Badri Vishal, et al.
Published: (2025)
by: Kasuba, Badri Vishal, et al.
Published: (2025)
INSITE: labelling medical images using submodular functions and semi-supervised data programming
by: Gautam, Akshat, et al.
Published: (2024)
by: Gautam, Akshat, et al.
Published: (2024)
Improving Selective Classification with Pairwise Queries for Binary Classification
by: Vardhan, Harsh, et al.
Published: (2026)
by: Vardhan, Harsh, et al.
Published: (2026)
Struct-X: Enhancing Large Language Models Reasoning with Structured Data
by: Tan, Xiaoyu, et al.
Published: (2024)
by: Tan, Xiaoyu, et al.
Published: (2024)
Centered Masking for Language-Image Pre-Training
by: Liang, Mingliang, et al.
Published: (2024)
by: Liang, Mingliang, et al.
Published: (2024)
Masked Structural Growth for 2x Faster Language Model Pre-training
by: Yao, Yiqun, et al.
Published: (2023)
by: Yao, Yiqun, et al.
Published: (2023)
Fed-SB: A Silver Bullet for Extreme Communication Efficiency and Performance in (Private) Federated LoRA Fine-Tuning
by: Singhal, Raghav, et al.
Published: (2025)
by: Singhal, Raghav, et al.
Published: (2025)
Analysing The Impact of Sequence Composition on Language Model Pre-Training
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Initialization using Update Approximation is a Silver Bullet for Extremely Efficient Low-Rank Fine-Tuning
by: Ponkshe, Kaustubh, et al.
Published: (2024)
by: Ponkshe, Kaustubh, et al.
Published: (2024)
StructLM: Towards Building Generalist Models for Structured Knowledge Grounding
by: Zhuang, Alex, et al.
Published: (2024)
by: Zhuang, Alex, et al.
Published: (2024)
StructLens: A Structural Lens for Language Models via Maximum Spanning Trees
by: Sakajo, Haruki, et al.
Published: (2026)
by: Sakajo, Haruki, et al.
Published: (2026)
Fisher Mask Nodes for Language Model Merging
by: K, Thennal D, et al.
Published: (2024)
by: K, Thennal D, et al.
Published: (2024)
KAUCUS: Knowledge Augmented User Simulators for Training Language Model Assistants
by: Dhole, Kaustubh D.
Published: (2024)
by: Dhole, Kaustubh D.
Published: (2024)
LEVOS: Leveraging Vocabulary Overlap with Sanskrit to Generate Technical Lexicons in Indian Languages
by: J, Karthika N, et al.
Published: (2024)
by: J, Karthika N, et al.
Published: (2024)
Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
by: Guo, Xu, et al.
Published: (2026)
by: Guo, Xu, et al.
Published: (2026)
StructEval: Deepen and Broaden Large Language Model Assessment via Structured Evaluation
by: Cao, Boxi, et al.
Published: (2024)
by: Cao, Boxi, et al.
Published: (2024)
On the Diversity of Synthetic Data and its Impact on Training Large Language Models
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Masked Diffusion Language Models with Frequency-Informed Training
by: Kosmopoulou, Despoina, et al.
Published: (2025)
by: Kosmopoulou, Despoina, et al.
Published: (2025)
Modular Techniques for Synthetic Long-Context Data Generation in Language Model Training and Evaluation
by: Subramanian, Seganrasan, et al.
Published: (2025)
by: Subramanian, Seganrasan, et al.
Published: (2025)
Long Context Pre-Training with Lighthouse Attention
by: Peng, Bowen, et al.
Published: (2026)
by: Peng, Bowen, et al.
Published: (2026)
Legal Documents Drafting with Fine-Tuned Pre-Trained Large Language Model
by: Lin, Chun-Hsien, et al.
Published: (2024)
by: Lin, Chun-Hsien, et al.
Published: (2024)
StructTest: Benchmarking LLMs' Reasoning through Compositional Structured Outputs
by: Chen, Hailin, et al.
Published: (2024)
by: Chen, Hailin, et al.
Published: (2024)
StructKV: Preserving the Structural Skeleton for Scalable Long-Context Inference
by: Chen, Zhirui, et al.
Published: (2026)
by: Chen, Zhirui, et al.
Published: (2026)
HiStruct+: Improving Extractive Text Summarization with Hierarchical Structure Information
by: Ruan, Qian, et al.
Published: (2022)
by: Ruan, Qian, et al.
Published: (2022)
ProtStructQA: A Denotation Threshold in Protein Structural Reasoning
by: Mandiga, Aravind, et al.
Published: (2026)
by: Mandiga, Aravind, et al.
Published: (2026)
A Code Comprehension Benchmark for Large Language Models for Code
by: Havare, Jayant, et al.
Published: (2025)
by: Havare, Jayant, et al.
Published: (2025)
Development of Pre-Trained Transformer-based Models for the Nepali Language
by: Thapa, Prajwal, et al.
Published: (2024)
by: Thapa, Prajwal, et al.
Published: (2024)
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
by: Maheshwari, Ayush, et al.
Published: (2022)
by: Maheshwari, Ayush, et al.
Published: (2022)
Power Mechanism: Private Tabular Representation Release for Model Agnostic Consumption
by: Vepakomma, Praneeth, et al.
Published: (2025)
by: Vepakomma, Praneeth, et al.
Published: (2025)
StructCoh: Structured Contrastive Learning for Context-Aware Text Semantic Matching
by: Xue, Chao, et al.
Published: (2025)
by: Xue, Chao, et al.
Published: (2025)
MHQA: A Diverse, Knowledge Intensive Mental Health Question Answering Challenge for Language Models
by: Racha, Suraj, et al.
Published: (2025)
by: Racha, Suraj, et al.
Published: (2025)
Characterizing the Behavior of Training Mamba-based State Space Models on GPUs
by: Baruah, Trinayan, et al.
Published: (2025)
by: Baruah, Trinayan, et al.
Published: (2025)
Similar Items
-
GenSco: Can Question Decomposition based Passage Alignment improve Question Answering?
by: Fazili, Barah, et al.
Published: (2024) -
GUIDEQ: Framework for Guided Questioning for progressive informational collection and classification
by: Mishra, Priya, et al.
Published: (2024) -
FedEx-LoRA: Exact Aggregation for Federated and Efficient Fine-Tuning of Foundation Models
by: Singhal, Raghav, et al.
Published: (2024) -
SPRINT: Script-agnostic Structure Recognition in Tables
by: Kudale, Dhruv, et al.
Published: (2025) -
A Lightweight Method to Disrupt Memorized Sequences in LLM
by: Prashant, Parjanya Prajakta, et al.
Published: (2025)