--- id: swebench version: "4.1.0" license: MIT License Copyright (c) 2023 Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, Karthik R Narasimhan Permission is hereby granted, free of charge, to any person… (full text in the JSON record) license_treatment: permissive maintenance: active --- # swebench — The official SWE-bench package - a benchmark for evaluating LMs on software engineering License: permissive · Maintenance: active · Popularity: top 1,000 on PyPI ## Install pip install swebench uv add swebench poetry add swebench ## Description

Kawi the SWE-Llama

Read the Docs ]

日本語 | 中文简体 | 中文繁體

Build License

--- Code and data for the following works: * [ICLR 2025] SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? * [ICLR 2024 Oral] SWE-bench: Can Language Models Resolve... ## AI interpretation — verify before relying SWE-bench is a benchmark for evaluating language models on real-world GitHub software issues, where models generate patches to resolve described problems in codebases. Verdict: Actively maintained, well-resourced benchmark with no known vulnerabilities and permissive licensing. Suitable for evaluating LM code-generation capabilities, but evaluation is resource-intensive and requires Docker infrastructure setup. [View on SkillFed](https://skillfed.io/packages/swebench) · [View on PyPI](https://pypi.org/project/swebench/)