HotpotQA
Emerging39papers using it
2023first seen
Dataset Card for BEIR Benchmark hotpotqa is one of the datasets from the Question Answering task within BEIR, measuring Wikipedia article retrieval for a given multi-hop query. Dataset Summary BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks. Fact-checking: FEV
Papers using HotpotQA (29)
- Learning Query-Aware Budget-Tier Routing for Runtime Agent MemoryTool-Schema Compression Enables Agentic RAG Under Constrained Context BudgetsKnowAgent: Knowledge-Augmented Planning for LLM-Based AgentsCodeAgents: A Token-Efficient Framework for Codified Multi-Agent Reasoning in LLMsBridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic SearchTrack, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language AgentsMemPro: Agentic Memory Systems as Evolvable ProgramsCascading Hallucination in Agentic RAG: The CHARM Framework for Detection and MitigationAdaMEM: Test-Time Adaptive Memory for Language AgentsSemantic Early-Stopping for Iterative LLM Agent LoopsWhen Latent Agents Lie: KV-Cache Integrity in Multi-Agent LLM CollaborationContrastive Reflection for Iterative Prompt OptimizationZEBRA: Zero-shot Budgeted Resource Allocation for LLM OrchestrationParallel Context Compaction for Long-Horizon LLM Agent ServingProper Scoring Rules for Agentic Uncertainty QuantificationRetrieval as Reasoning: Self-Evolving Agent-Native Retrieval via LLM-WikiStepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement LearningPrompt Codebooks: Discrete Compositional Optimization for Language Model Instruction RefinementCritic-R: Improving Agentic Search using Instruction-tuned Retrievers with Natural Language Introspective FeedbackAnswer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic SystemsGRASP: Graph Agentic Search over Propositions for Multi-hop Question AnsweringDecoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM AgentsScaling Multi-agent Systems: A Smart Middleware for Improving Agent InteractionsMemSkill: Learning and Evolving Memory Skills for Self-Evolving AgentsPseudoAct: Leveraging Pseudocode Synthesis for Flexible Planning and Action Control in Large Language Model AgentsMAR:Multi-Agent Reflexion Improves Reasoning Abilities in LLMsMission Impossible: Feedback-Guided Dynamic Interactive Planning for Improving Reasoning on LLMsDebFlow: Automating Agent Creation via Agent DebateSmurfs: Multi-agent System Using Context-efficient DFSDT For Tool Planning