$FastDoc$: Domain-Specific Fast Continual Pre-training Technique using Document-Level Metadata and Taxonomy
Fuente:
arXiv
Saved in:
| Main Authors: | Nandy, Abhilash, Kapadnis, Manav Nitin, Patnaik, Sohan, Butala, Yash Parag, Goyal, Pawan, Ganguly, Niloy |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SERPENT-VLM : Self-Refining Radiology Report Generation Using Vision Language Models
by: Kapadnis, Manav Nitin, et al.
Published: (2024)
by: Kapadnis, Manav Nitin, et al.
Published: (2024)
Order-Based Pre-training Strategies for Procedural Text Understanding
by: Nandy, Abhilash, et al.
Published: (2024)
by: Nandy, Abhilash, et al.
Published: (2024)
YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models
by: Nandy, Abhilash, et al.
Published: (2024)
by: Nandy, Abhilash, et al.
Published: (2024)
CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review
by: Kapadnis, Manav Nitin, et al.
Published: (2025)
by: Kapadnis, Manav Nitin, et al.
Published: (2025)
Label-semantics Aware Generative Approach for Domain-Agnostic Multilabel Classification
by: Khatuya, Subhendu, et al.
Published: (2025)
by: Khatuya, Subhendu, et al.
Published: (2025)
On The Persona-based Summarization of Domain-Specific Documents
by: Mullick, Ankan, et al.
Published: (2024)
by: Mullick, Ankan, et al.
Published: (2024)
ChartEditBench: Evaluating Grounded Multi-Turn Chart Editing in Multimodal Language Models
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative Retrieval
by: Khatuya, Subhendu, et al.
Published: (2025)
by: Khatuya, Subhendu, et al.
Published: (2025)
REFINE-AF: A Task-Agnostic Framework to Align Language Models via Self-Generated Instructions using Reinforcement Learning from Automated Feedback
by: Roy, Aniruddha, et al.
Published: (2025)
by: Roy, Aniruddha, et al.
Published: (2025)
A Pointer Network-based Approach for Joint Extraction and Detection of Multi-Label Multi-Class Intents
by: Mullick, Ankan, et al.
Published: (2024)
by: Mullick, Ankan, et al.
Published: (2024)
Known Intents, New Combinations: Clause-Factorized Decoding for Compositional Multi-Intent Detection
by: Nandy, Abhilash
Published: (2026)
by: Nandy, Abhilash
Published: (2026)
IDALC: A Semi-Supervised Framework for Intent Detection and Active Learning based Correction
by: Mullick, Ankan, et al.
Published: (2025)
by: Mullick, Ankan, et al.
Published: (2025)
Latent Diffusion Pretraining for Crystal Property Prediction
by: Mukherjee, Shrimon, et al.
Published: (2026)
by: Mukherjee, Shrimon, et al.
Published: (2026)
Instruction-Guided Bullet Point Summarization of Long Financial Earnings Call Transcripts
by: Khatuya, Subhendu, et al.
Published: (2024)
by: Khatuya, Subhendu, et al.
Published: (2024)
Efficient Continual Pre-training of LLMs for Low-resource Languages
by: Nag, Arijit, et al.
Published: (2024)
by: Nag, Arijit, et al.
Published: (2024)
REVEAL -- Reasoning and Evaluation of Visual Evidence through Aligned Language
by: Praharaj, Ipsita, et al.
Published: (2025)
by: Praharaj, Ipsita, et al.
Published: (2025)
How Robust are the Tabular QA Models for Scientific Tables? A Study using Customized Dataset
by: Ghosh, Akash, et al.
Published: (2024)
by: Ghosh, Akash, et al.
Published: (2024)
MEDVOC: Vocabulary Adaptation for Fine-tuning Pre-trained Language Models on Medical Text Summarization
by: Balde, Gunjan, et al.
Published: (2024)
by: Balde, Gunjan, et al.
Published: (2024)
LLM Meets Diffusion: A Hybrid Framework for Crystal Material Generation
by: Khastagir, Subhojyoti, et al.
Published: (2025)
by: Khastagir, Subhojyoti, et al.
Published: (2025)
Periodic Materials Generation using Text-Guided Joint Diffusion Model
by: Das, Kishalay, et al.
Published: (2025)
by: Das, Kishalay, et al.
Published: (2025)
DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization
by: Chau, Phan Phuong Mai, et al.
Published: (2024)
by: Chau, Phan Phuong Mai, et al.
Published: (2024)
VoiceDoc: A Voice-Activated Intelligent Document Assistant Using Advanced RAG Technology
by: Patel, Manav
Published: (2026)
by: Patel, Manav
Published: (2026)
PBEBench: A Multi-Step Programming by Examples Reasoning Benchmark inspired by Historical Linguistics
by: Naik, Atharva, et al.
Published: (2025)
by: Naik, Atharva, et al.
Published: (2025)
Long Dialog Summarization: An Analysis
by: Mullick, Ankan, et al.
Published: (2024)
by: Mullick, Ankan, et al.
Published: (2024)
Fine-Grained Classification: Connecting Metadata via Cross-Contrastive Pre-Training
by: Mamtani, Sumit, et al.
Published: (2025)
by: Mamtani, Sumit, et al.
Published: (2025)
Read the Docs Before Rewriting: Equip Rewriter with Domain Knowledge via Continual Pre-training
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Parameter-Efficient Instruction Tuning of Large Language Models For Extreme Financial Numeral Labelling
by: Khatuya, Subhendu, et al.
Published: (2024)
by: Khatuya, Subhendu, et al.
Published: (2024)
Leveraging Self-Attention for Input-Dependent Soft Prompting in LLMs
by: Muppidi, Ananth, et al.
Published: (2025)
by: Muppidi, Ananth, et al.
Published: (2025)
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
by: Kapoor, Raghav, et al.
Published: (2024)
by: Kapoor, Raghav, et al.
Published: (2024)
DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-Completion
by: Mehta, Rahul, et al.
Published: (2026)
by: Mehta, Rahul, et al.
Published: (2026)
Pre-training Graph Neural Networks with Structural Fingerprints for Materials Discovery
by: Jia, Shuyi, et al.
Published: (2025)
by: Jia, Shuyi, et al.
Published: (2025)
$\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
by: Das, Trishanu, et al.
Published: (2025)
by: Das, Trishanu, et al.
Published: (2025)
Metadata Conditioning Accelerates Language Model Pre-training
by: Gao, Tianyu, et al.
Published: (2025)
by: Gao, Tianyu, et al.
Published: (2025)
AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models
by: Patnaik, Sohan, et al.
Published: (2025)
by: Patnaik, Sohan, et al.
Published: (2025)
It Helps to Take a Second Opinion: Teaching Smaller LLMs to Deliberate Mutually via Selective Rationale Optimisation
by: Patnaik, Sohan, et al.
Published: (2025)
by: Patnaik, Sohan, et al.
Published: (2025)
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning
by: Patnaik, Sohan, et al.
Published: (2025)
by: Patnaik, Sohan, et al.
Published: (2025)
DocMamba: Efficient Document Pre-training with State Space Model
by: Hu, Pengfei, et al.
Published: (2024)
by: Hu, Pengfei, et al.
Published: (2024)
Leveraging the Power of LLMs: A Fine-Tuning Approach for High-Quality Aspect-Based Summarization
by: Mullick, Ankan, et al.
Published: (2024)
by: Mullick, Ankan, et al.
Published: (2024)
Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documents
by: Mullick, Ankan, et al.
Published: (2025)
by: Mullick, Ankan, et al.
Published: (2025)
Leveraging Large Language Models for Predictive Analysis of Human Misery
by: Seal, Bishanka, et al.
Published: (2025)
by: Seal, Bishanka, et al.
Published: (2025)
Similar Items
-
SERPENT-VLM : Self-Refining Radiology Report Generation Using Vision Language Models
by: Kapadnis, Manav Nitin, et al.
Published: (2024) -
Order-Based Pre-training Strategies for Procedural Text Understanding
by: Nandy, Abhilash, et al.
Published: (2024) -
YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models
by: Nandy, Abhilash, et al.
Published: (2024) -
CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review
by: Kapadnis, Manav Nitin, et al.
Published: (2025) -
Label-semantics Aware Generative Approach for Domain-Agnostic Multilabel Classification
by: Khatuya, Subhendu, et al.
Published: (2025)