How Does India Cook Biryani?
Fuente:
arXiv
Guardado en:
| Autores principales: | Goel, Shubham, S, Farzana, Rishi, C V, Arun, Aditya, Jawahar, C V |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education
por: M, Megha Mariam K., et al.
Publicado: (2026)
por: M, Megha Mariam K., et al.
Publicado: (2026)
Attend to what I say: Highlighting relevant content on slides
por: M, Megha Mariam K, et al.
Publicado: (2026)
por: M, Megha Mariam K, et al.
Publicado: (2026)
IndicSTR12: A Dataset for Indic Scene Text Recognition
por: Lunia, Harsh, et al.
Publicado: (2024)
por: Lunia, Harsh, et al.
Publicado: (2024)
Advancing Question Answering on Handwritten Documents: A State-of-the-Art Recognition-Based Model for HW-SQuAD
por: Pal, Aniket, et al.
Publicado: (2024)
por: Pal, Aniket, et al.
Publicado: (2024)
Source-free Video Domain Adaptation by Learning from Noisy Labels
por: Dasgupta, Avijit, et al.
Publicado: (2023)
por: Dasgupta, Avijit, et al.
Publicado: (2023)
Prompt2LVideos: Exploring Prompts for Understanding Long-Form Multimodal Videos
por: Jahagirdar, Soumya Shamarao, et al.
Publicado: (2025)
por: Jahagirdar, Soumya Shamarao, et al.
Publicado: (2025)
Real-Time Human Reconstruction and Animation using Feed-Forward Gaussian Splatting
por: Chatterjee, Devdoot, et al.
Publicado: (2026)
por: Chatterjee, Devdoot, et al.
Publicado: (2026)
EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer
por: Monga, Munish, et al.
Publicado: (2026)
por: Monga, Munish, et al.
Publicado: (2026)
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
por: Pal, Aniket, et al.
Publicado: (2025)
por: Pal, Aniket, et al.
Publicado: (2025)
PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction
por: Mishra, Naman, et al.
Publicado: (2026)
por: Mishra, Naman, et al.
Publicado: (2026)
Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media
por: M, Megha Mariam K., et al.
Publicado: (2026)
por: M, Megha Mariam K., et al.
Publicado: (2026)
Face Time Traveller : Travel Through Ages Without Losing Identity
por: Kar, Purbayan, et al.
Publicado: (2026)
por: Kar, Purbayan, et al.
Publicado: (2026)
Reading Between the Lanes: Text VideoQA on the Road
por: Tom, George, et al.
Publicado: (2023)
por: Tom, George, et al.
Publicado: (2023)
Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach
por: Khindkar, Vaishnavi, et al.
Publicado: (2024)
por: Khindkar, Vaishnavi, et al.
Publicado: (2024)
How Does Bilateral Ear Symmetry Affect Deep Ear Features?
por: Ozturk, Kagan, et al.
Publicado: (2025)
por: Ozturk, Kagan, et al.
Publicado: (2025)
DriveSafe: A Framework for Risk Detection and Safety Suggestions in Driving Scenarios
por: Artham, Sainithin, et al.
Publicado: (2026)
por: Artham, Sainithin, et al.
Publicado: (2026)
Pedestrian Intention and Trajectory Prediction in Unstructured Traffic Using IDD-PeD
por: Bokkasam, Ruthvik, et al.
Publicado: (2025)
por: Bokkasam, Ruthvik, et al.
Publicado: (2025)
AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval
por: Maniyar, Suyash, et al.
Publicado: (2025)
por: Maniyar, Suyash, et al.
Publicado: (2025)
IDD-X: A Multi-View Dataset for Ego-relative Important Object Localization and Explanation in Dense and Unstructured Traffic
por: Parikh, Chirag, et al.
Publicado: (2024)
por: Parikh, Chirag, et al.
Publicado: (2024)
Multiple Instance Learning for Glioma Diagnosis using Hematoxylin and Eosin Whole Slide Images: An Indian Cohort Study
por: Chauhan, Ekansh, et al.
Publicado: (2024)
por: Chauhan, Ekansh, et al.
Publicado: (2024)
The More You See in 2D, the More You Perceive in 3D
por: Han, Xinyang, et al.
Publicado: (2024)
por: Han, Xinyang, et al.
Publicado: (2024)
Towards Accurate Lip-to-Speech Synthesis in-the-Wild
por: Hegde, Sindhu, et al.
Publicado: (2024)
por: Hegde, Sindhu, et al.
Publicado: (2024)
Chain-of-Cooking:Cooking Process Visualization via Bidirectional Chain-of-Thought Guidance
por: Xu, Mengling, et al.
Publicado: (2025)
por: Xu, Mengling, et al.
Publicado: (2025)
Towards Deployable OCR models for Indic languages
por: Mathew, Minesh, et al.
Publicado: (2022)
por: Mathew, Minesh, et al.
Publicado: (2022)
A Dataset for Semantic Segmentation in the Presence of Unknowns
por: Laskar, Zakaria, et al.
Publicado: (2025)
por: Laskar, Zakaria, et al.
Publicado: (2025)
Towards Safer and Understandable Driver Intention Prediction
por: Karuppasamy, Mukilan, et al.
Publicado: (2025)
por: Karuppasamy, Mukilan, et al.
Publicado: (2025)
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
por: Plizzari, Chiara, et al.
Publicado: (2024)
por: Plizzari, Chiara, et al.
Publicado: (2024)
Cross-Domain Evaluation of Few-Shot Classification Models: Natural Images vs. Histopathological Images
por: Sekhar, Ardhendu, et al.
Publicado: (2024)
por: Sekhar, Ardhendu, et al.
Publicado: (2024)
Hide and Seek: How Does Watermarking Impact Face Recognition?
por: Yao, Yuguang, et al.
Publicado: (2024)
por: Yao, Yuguang, et al.
Publicado: (2024)
Utilizing Multi-Agent Reinforcement Learning with Encoder-Decoder Architecture Agents to Identify Optimal Resection Location in Glioblastoma Multiforme Patients
por: Arun, Krishna, et al.
Publicado: (2025)
por: Arun, Krishna, et al.
Publicado: (2025)
Designing Production-Scale OCR for India: Multilingual and Domain-Specific Systems
por: Faraz, Ali, et al.
Publicado: (2026)
por: Faraz, Ali, et al.
Publicado: (2026)
Curvature Informed Furthest Point Sampling
por: Bhardwaj, Shubham, et al.
Publicado: (2024)
por: Bhardwaj, Shubham, et al.
Publicado: (2024)
PRECISe : Prototype-Reservation for Explainable Classification under Imbalanced and Scarce-Data Settings
por: Ganatra, Vaibhav, et al.
Publicado: (2024)
por: Ganatra, Vaibhav, et al.
Publicado: (2024)
CookingDiffusion: Cooking Procedural Image Generation with Stable Diffusion
por: Wang, Yuan, et al.
Publicado: (2025)
por: Wang, Yuan, et al.
Publicado: (2025)
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution
por: Liang, Dingkang, et al.
Publicado: (2025)
por: Liang, Dingkang, et al.
Publicado: (2025)
VisualChef: Generating Visual Aids in Cooking via Mask Inpainting
por: Kuzyk, Oleh, et al.
Publicado: (2025)
por: Kuzyk, Oleh, et al.
Publicado: (2025)
How Does Audio Influence Visual Attention in Omnidirectional Videos? Database and Model
por: Zhu, Yuxin, et al.
Publicado: (2024)
por: Zhu, Yuxin, et al.
Publicado: (2024)
Real-Time Cooked Food Image Synthesis and Visual Cooking Progress Monitoring on Edge Devices
por: Gupta, Jigyasa, et al.
Publicado: (2025)
por: Gupta, Jigyasa, et al.
Publicado: (2025)
Novel View Synthesis using DDIM Inversion
por: Singh, Sehajdeep, et al.
Publicado: (2025)
por: Singh, Sehajdeep, et al.
Publicado: (2025)
V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos
por: Yue, Chengkun, et al.
Publicado: (2026)
por: Yue, Chengkun, et al.
Publicado: (2026)
Ejemplares similares
-
PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education
por: M, Megha Mariam K., et al.
Publicado: (2026) -
Attend to what I say: Highlighting relevant content on slides
por: M, Megha Mariam K, et al.
Publicado: (2026) -
IndicSTR12: A Dataset for Indic Scene Text Recognition
por: Lunia, Harsh, et al.
Publicado: (2024) -
Advancing Question Answering on Handwritten Documents: A State-of-the-Art Recognition-Based Model for HW-SQuAD
por: Pal, Aniket, et al.
Publicado: (2024) -
Source-free Video Domain Adaptation by Learning from Noisy Labels
por: Dasgupta, Avijit, et al.
Publicado: (2023)