Locked learning resources

Join us and get access to thousands of tutorials and a community of expert Pythonistas.

Unlock This Lesson

Locked learning resources

This lesson is for members only. Join us and get access to thousands of tutorials and a community of expert Pythonistas.

Unlock This Lesson

The Real Python AI Benchmark: Which Model Writes Good Python?

Every week there’s a new AI model and a new set of benchmark charts, and none of them tell you whether the thing writes good Python code. That’s why Real Python built an AI benchmark for Python developers. In this team call, Dan Bader and Martin Breuss walk Philipp Acsany through it.

This is a starting point rather than a finished league table, including the tests that failed to separate the models at all. Tell us in the comments which models you want us to add next.

Resources mentioned in this lesson:

A cartoon python wearing glasses points to a benchmark chart beside racks of testing equipment, gauges, an AI chip, and a Python logo.

The Real Python AI Benchmark

Compare AI models on everyday Python tasks: modern code, release knowledge, focused edits, and made-up functions. Start with a python reading a book.

00:00 It almost feels like every week there is a new AI model and literally every week there is a new AI model and when a new AI model comes out there are all these benchmarks from Anthropic and OpenAI and so on.

00:12 But how relevant is that really for you as a Python developer? So we put our heads together or more specifically Dan and Martin put their heads together to come up with a real Python benchmark.

00:23 Hi Dan, welcome to this call. I haven’t seen you guys in in ages. Hi Martin. Hi, good to see you. So you both were quite active over the last couple of weeks and what came out is a real Python benchmark for AI models.

00:42 So maybe you can tell our viewers a little bit what this is, why you came up with this idea in the first place Dan and then maybe Martin you can talk a bit about the technicalities of it.

00:52 Like you said, you know there’s so much news happening in AI and like this is the topic I think for anyone in technology, anyone who has a career in software development.

01:01 We have to keep up with that stuff and it’s a very exciting time and it’s hard to make sense of all the noise. So Martin and I, we had a couple of conversations about this problem and Philipp you were also part of many of these.

01:14 We want to help learners and readers and students that use real Python to sort of cut through the noise and really fully participate in this conversation.

01:24 Yeah, I want to hand it over to you Martin. Maybe you can talk about what the kind of the goal of the benchmark was or kind of what the elements are that go into it so that people understand what they’re looking at on their screen right there.

01:36 Right, there’s like lots of professional benchmarks out there that give you some sense of models capabilities but it doesn’t necessarily translate into your day-to-day as a Python programmer.

01:47 Like should you really try like work with this different model now just because it came out and what’s actually going to be the impact on the work that you do.

01:55 I’ve been trying to figure out what are the different types of work that a Python developer does and try to you know set up a couple of benchmarks related to that and then compare the different models so that we could give a good idea of how relevant is this for you in a real Python developer situation.

02:12 And it was hard to get those benchmarks right. I was running it on frontier models and no matter what I came up with, most of them were just acing it and there was no differentiator essentially that would be relevant or interesting to talk about.

02:27 So I kept banging my head against this for a while and then eventually, also inspired by Simon Willison’s Pelican on the Bicycle, I was thinking about what could be a Pythonic version of this that would still give readers like a sense of how good is this with Python and also be like fun and interesting to look at right.

02:47 Eventually, I decided on Python, right, reading a book because we’re a learning platform. It’s all about learning the mascot of Python obviously. I chose to use the turtle module because I feel like there’s something inherently Pythonic about the turtle module that probably a lot of developers that have started with Python a while ago have eventually encountered doing something with turtle.

03:08 So what you’re seeing here, because I think this is this is really cool actually, is like so Simon, obviously we’re standing on the shoulders of giants there, right, like we love Simon and his Pelicans.

03:18 I mean, it’s just such a fun little benchmark. So like our twist on this is essentially instead of asking the model to draw an SVG file, we’re basically asking it to run the turtle module in Python.

03:30 This is literally the prompt for this, right? Write a Python turtle program that draws a Python reading a book like this is the prompt for other models for this one part of the benchmark.

03:41 I’ll talk about the rest too in just a second. But a fun thing that this gives us also is that the way that I’ve implemented it here is that you can actually watch it, draw it by just taking the specifics.

03:51 There’s some weird ones. So if you go further down and you look at some older models, like, I don’t know, GPT-4o, like, what is this? Right? Like, it’s not hiding the turtle.

04:02 You can see the turtle still moving. And then I guess it like messed up where the book is going to go. I believe this is like the text in the book. And look at the release date, right?

04:13 This is like May 2024. And so just the jump from this, if you look at this, and then you look at Fable or Astra is amazing. That’s also an important bit about this benchmark, because we’ll go more into some details of other benchmarks there as well.

04:31 But we as humans, we are visual, and we like to see something. And that’s what I like about drawing benchmarks is like, you get immediately a feeling. Like, it doesn’t really say much about how well this model is with coding, but it kind of gives you a first impression.

04:46 And it’s like a bit silly sometimes, and you can make out of it. But I think it’s like also impressive then to see similar versions of models advancing in that.

04:55 So that’s like something we aspire to over time. Like once we have more and more models in our benchmark overview, you can also see like how it changed over time.

05:03 And even like you were saying, like older models, but some of them are just like a few weeks old, and you could already see that there is a bit of a jump there.

05:13 And I like this turtle twist. So it’s kind of like a little Python program that every model worked. And you were saying, Martin, that was one of the things that was a bit difficult for you at the beginning, because if you just throw a coding task at a model, it just like performs the coding task and like it did well.

05:30 So what were other benchmarks that you came up with then in order to really see if a model is really good at writing Python code? One of the things that I encountered frequently when I use AI to write some Python code is from __future__ import annotations for typing, right?

05:48 Like this is something that a lot of AI models still produce. And it’s just not necessary anymore, because 3.14, like you don’t need any of that anymore. You can just use the list directly to write your type annotations, right?

06:00 And so I was thinking about what’s the knowledge cutoff in terms of Python, that would be something that I would like to know before I start working with one of these new models.

06:11 And out of that kind of came this like writes Python like it’s 2021, and then have a specific version of Python, like so that was the version of Python that was out in that year.

06:23 There’s a little program that it is asked to write just against the specs of a program to write that has like a lot of things in there, where it could potentially trip up like, it’s not really tripping up, it’s still writing valid code, right?

06:37 But it’s writing code based on a Python of the past, if it doesn’t know about the new development of the language, essentially. That’s good to know, as a Python developer, where you might steer the model a bit more by even like if you’re having your agentic harness to already have your environment set up in a way to address a certain Python version.

06:55 But if you just have a Python question about something, you really want to make sure that it’s best practice of the current Python, right? There’s also this question of what’s the newest Python that it knows?

07:05 And that’s just a straight up question. So the benchmark also asks this question. So what’s the newest Python that you know of basically? And that’s just reply with just a number basically, right?

07:15 And this is the output that we get here. No web search, no external research, right? It’s like, what is like,

07:21 yeah, yeah. And so it’s also interesting to see that this can be different, right? So it knows off. It’s like, oh, yeah, I know 3.13. But then it still uses idioms where it could reach for something more modern.

07:32 When writing code, it still reaches for something of a previous Python version. Can I write that the newest Python that Astra knows is actually somewhat behind, like a Fable seems to know about 3.14.

07:44 And then Astra doesn’t is really aware that that’s

07:48 happened, which is that’s interesting. It’s actually pretty annoying if you hit an edge case where that’s relevant. Like, for example, I had this experience where, you know, new version of Ruff came out.

07:58 And this was, I don’t know, at this point, a couple weeks, months ago. And I was like, okay, cool, great. I’ll upgrade it, right? I have like reformat all of my code.

08:05 And I started using like a 3.14 feature, I think was the one where you can drop the parentheses when you’re catching multiple exception classes in one except statement.

08:14 And so that was a new thing in 3.14, if I remember correctly. And so Ruff now wants this, right? So it’ll like reformat and like flag it every time that’s not it’s not the parentheses are still there.

08:26 And then for a while, I was fighting the model, which was, I think it was either Opus 5 or like the original Fable. I forgot which one got it right. But for sure, Fable 5.1 now is aware of that.

08:36 And so you have this like back and forth between like the model changing it back and then the linter triggering and it’d be like, this can’t be right. Like this is not valid Python.

08:44 And then it would do like a web search. And then be like, okay, now it is valid. Okay, great. Now we can continue. And so we do that every single session, because it was like, this is broken Python, the infrastructure is broken, we have to update the linter and would do like burn all these tokens on these like site quests.

08:59 And so if you have a model, I mean, that knows that 3.14 exists, it’ll it’ll not do that. And so that’s that’s a huge win for for usability. And you don’t have to fix it manually.

09:09 The price difference is something I think that really stands out. I mean, just looking at the cost there for these runs, that’s also something that’s relevant.

09:16 I feel I mean, obviously, that’s relevant as a Python developer, right? Like if you’re if what are you actually paying for the intelligence that you get from this different models.

09:24 So I mean, these are like relatively small tasks. They’re not like ingesting giant amounts of code, etc. But they’re all they all get the same tasks. So you still get a comparison, obviously, what what this costs are what you can expect to get.

09:36 And then, yeah, so so here, for example, if you compare Fable here with Astra, it’s like significant difference of like Claude Fable being just a lot more expensive for what what it used to get to this same results like to run through this whole benchmark tasks.

09:52 Like one thing that I’ve also noticed is like this lines touched for a tiny edit, which also came from, for me, from one of the annoyances that I hear a lot and that people experience when they when they code with models is like, oh, I just gave it a small task and refactored my whole code base.

10:06 I mean, you’re probably on a too old model if that’s what you’re getting. But just like, how much does it actually do versus how much does it need to do? Right.

10:13 So we have this article standard, where you can take a look at what actually was changed. So here we have the file change diff of those of that little command line script.

10:24 Right. And here you can see like this is, for example, it adds the help text here and it formats it on separate lines versus not doing that. And that’s I think that’s literally the whole difference to the optimal solution.

10:38 So not really differentiated that much. But I would like to find something in that sense. But maybe this is also the path I’ve been going down before with with previous iterations of this benchmark.

10:50 I was like, I think this is really important to get a differentiator based on how many of these bugs does it find. Right. But then it wasn’t actually a differentiator.

10:59 So that’s, I guess, the tricky bits of designing a good benchmark that actually is helpful for usage. Right. But the nice thing about benchmarks is and that’s why we have it in these different categories is like, yeah, maybe you don’t care about like how well can it draw a Python reading a book.

11:16 It’s more about what’s the newest Python and there of Muse Spark, for example, with 3.14.1. It’s like very recent. So I think that is that is nice. And the other part that you were saying is like true.

11:27 I mean, with with the benchmark, it’s it needs to have some kind of like differentiating factor. It’s not like, hey, all the models are are fine, but we want to find the things where they differ.

11:37 And it could be like that now the models that you put to test are like fine with this lines touched with this tiny edit benchmark. But part of benchmark is also if there is a new model coming out, maybe it like adds doc strings to all the functions by default.

11:54 And then we see like that there maybe was a digression there. So I think that’s interesting about benchmarks in general. This goes both ways. I remember this was one of the older, I think, Sonnet models that all of a sudden started like adding comments to every everywhere in your code base, like inline comments.

12:12 I think that’s something we’d catch immediately with this type of even very sort of vibe check type of type of benchmark. And that’s a very useful signal.

12:21 Yeah. Yeah. I mean, this is basically like a starting point for real Python where we’re really interested in this technology. And I think it’s so impactful for any developer out there.

12:31 And so, you know, we want to have like a skeleton. We can hang more of these types of insights and sort of research that we’re doing for for our learners that use real Python.

12:41 And so I think this is just a great way to kind of organize it. Right. And it’s very visual and you can look at it and be like, OK, I just get a sense of, you know, maybe this new Muse Spark model quality wise, like it seems to just have a different kind of anatomical understanding of what it makes look like compared to something like Astra.

12:58 Yeah. So we really want to expand this and grow this. And obviously, also, if you have thoughts, you know, if you’re watching this right now, OK, you should cover this or here’s a model I really care about.

13:06 I don’t see it on the list. Let us know in the comments. And we’re very keen to add this this kind of coverage, just excited about what’s happening in the space there.

13:15 Yeah. Play with it. Take a look. Like we have two of those articles you can click through so far. We’ll add more. But you can look at Claude Fable and you can look at GPT-6 Astra and take a look at those.

13:26 You also have all the like the code that it produced in there. Right. So you can take a look at the checkpoints so you can learn more about how to what it actually checks for.

13:37 And you can also get you can also get the snake with that turtle

13:42 program and the script and run it on your computer if you want to try it out. So, yeah, go ahead. Take a look and play around with this, I would say.

13:53 Yeah. Thank you so much to give a first impression. So like you both mentioned, that is basically a starting point. So there will more models coming to the benchmark and we’re also looking for feedback.

14:06 Like what benchmarks are you interested in? What should we put the models to test? And then soon you will see it on this new benchmark that is around the corner.

14:17 So, yeah. Then, Martin, thanks for joining. One thing I really like about this benchmark is that you as a Python developer, you can actually dive in and see more about the benchmarks.

14:27 And that’s something I always missed with like those big benchmarks. It’s like, yeah, you see those numbers and like these big tests. But for you as a Python developer, you sometimes care more about the things that affect you in your daily life.

14:40 And with the articles that are behind those benchmarks, you can really dive in, see the code that the models write, and get a good first impression. So let us know what you think and see you around on realpython.com.

Become a Member to join the conversation.