What we saw the first time we used Cursor
In 2023 we tried a new code editor with GPT-3.5 wired into it. By the end of the afternoon it was clear that the cost of building software had changed, and that a studio built around that fact would have an advantage for a while.
Update, August 2026: SpaceX has since bought the company that makes Cursor for $60 billion. We have written a follow-up on what we got right and what we missed. The article below is as we published it.
Cursor came out in March 2023. Four MIT graduates had taken Visual Studio Code, the editor most developers already used, and rebuilt it so that a language model sat inside it with your code in front of it. At the time the models on offer were GPT-3.5 and GPT-4, and for anything you wanted to do quickly and often, GPT-3.5 was the one you used. It was fast and it was cheap.
We had been using generative AI for a while by then, mostly through a chat window: describe a problem, paste the answer back into the editor, fix what it got wrong. Useful, and a long way from changing how software gets made. Cursor was different in one way that mattered. The model could see the codebase. You described a change in plain English, in the file where the change belonged, and the code appeared there.
What we gave it
We gave it a real job, the kind of thing both of us had written a dozen times: read a messy export from one system, clean it, and load it into another, with the usual edge cases around dates, duplicates and missing fields.
GPT-3.5 got a fair amount of it wrong. It invented a library function that did not exist. It handled dates in a way that would have failed on the first British customer. It missed an edge case a junior engineer would also have missed.
Correcting it was faster than writing it. A senior engineer who knew what the code should do could read what came back, see where it was wrong, say so in a sentence, and have a corrected version in seconds. The work that took the time was no longer the typing. It was knowing what to ask for and knowing when the answer was wrong.
What that meant
We had both spent our careers on the other side of this. The cost of software was the cost of engineers' time, and the time went on implementation. Estimates, roadmaps, hiring plans and fundraising all followed from that. If the implementation cost fell by a large multiple, and the models were only going to get better at it, then most of what we knew about running a software company was about to be wrong.
Two things followed. The constraint on building software was going to move from the people who could type it to the people who could judge it. And a small team of experienced engineers, with tools like this, would for a period be able to do what used to take a large one, before the rest of the market caught up.
That period is what IRLY was set up to use. We founded the studio in 2024 around a small number of senior people using these tools every day, and around the early-stage companies who would benefit first from software becoming cheap to build well.
What we were not sure of
Plenty. GPT-3.5 was wrong often enough that we would not have let it near a customer's system without a person reading every change. Context windows were small, so anything larger than a few files needed careful handling. The tool itself was a few months old, made by a company with an $8 million seed round, and there was no guarantee it would be around in a year.
We made the bet anyway, because the direction was obvious even if the timing was not. The models would improve. The editors would improve. The cost of a working first version of a product would keep falling.
Share
Next
Why IRLYBuilding something?
We are a small studio in Leeds and we take on a limited number of projects at a time. Tell us what you are trying to do and we will tell you whether we are the right fit.
Start a conversation