Add community evaluation results for DEEP-SWE

#38
by nielsr HF Staff - opened

This PR adds community-provided evaluation results for the following benchmarks:

These results were extracted from the model card. This is based on the new evaluation results feature.

Note: This is an automated PR. Please review the evaluation results before merging.

Would it be possible to provide a breakdown of the DeepSWE results for DeepSeek-V4-Flash-0731? Specifically, I'm interested in knowing which tasks the model performs particularly well on and which tasks it struggles with. This would help the community better understand its strengths, weaknesses, and ideal use cases.

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment