← all papers · overview

Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi

Abstract

This research delved into GPT-4 and Kimi, two Large Language Models (LLMs), for systematic reviews. We evaluated their performance by comparing LLM-generated codes with human-generated codes from a peer-reviewed systematic review on assessment. Our findings suggested that the performance of LLMs fluctuates by data volume and question complexity for systematic reviews.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).