← all papers · overview

Evaluating Generative Language Models In Information Extraction As Subjective Question Correction

Abstract

Modern Large Language Models (LLMs) have showcased remarkable prowess in various tasks necessitating sophisticated cognitive behaviors. Nevertheless, a paradoxical performance discrepancy is observed, where these models underperform in seemingly elementary tasks like relation extraction and event extraction due to two issues in conventional evaluation. (1) The imprecision of existing evaluation me

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).