← all datasets

SWE-bench Multimodal

Emerging
4papers using it
2024first seen

The 'SWE-bench Multimodal' is a dataset/benchmark that contains multimodal data for evaluating automated program repair systems by requiring them to jointly reason over source code, textual issue descriptions, and visual artifacts like GUI screenshots.

Papers using SWE-bench Multimodal (4)

SWE-bench Multimodal dataset β€” papers, benchmarks & downloads Β· AI for Code