OpenAI FDE: “Code is very cheap to produce”
At the Shift conference in Zadar this September, we sat down with Luis Velasco, a Forward Deployed Engineer at OpenAI, who describes his role simply:
We help companies put AI to work in real environments.
We spoke with him about what it takes to move AI agents into production, how developers can make them reliable and secure, and which skills will matter most as AI changes software development.
Production starts with evals
As Luis hinted in the intro, Forward Deploy Engineers (FDE) help companies turn AI demos into real products. So we asked him what developers most often underestimate when they make that jump.

According to him, the real challenge starts when a demo reaches production. A demo can work well in controlled conditions, but production forces teams to understand exactly what works, what does not, and why failures happen.
That is where eval-driven (Evaluation-driven) development becomes important. He says OpenAI uses it extensively in its engagements to track and measure different workflows and tasks:
We have a scientific approach to measure the performance of the models doing a task. When, basically, you clear those benchmarks, those evals, you are good to go to production in a really good and safe way.
The deployment gap is the real challenge
We wanted to understand what makes an AI agent reliable enough for production, rather than something that only works well in a controlled environment.
Luis said it all comes down to “cracking the evals.” As he explained, today’s models already have impressive intelligence, but putting them to work in a company’s own environment requires giving them access to the company’s data, procedures, and tools:
You know, those three things, we kind of call the deployment gap, right? And this is what we’re trying to close as FD, basically.
Agents went from six-minute to 16-hour tasks
But what happens when an agent works on a task for hours rather than minutes? How should developers think about errors, memory, context, and human approval as those workflows become longer?
Luis answered with an example from his own experience. While checking the latest data from METR’s long time horizon benchmark, he was surprised by how quickly agents had improved. In just over three years, they went from handling tasks that would take a human around six minutes to tasks that would take around 16 hours.
That represents a massive jump, he said:
On one hand, you have the models getting better at long-term coherency and persistence, but on the other hand, you also have the harnesses improving a lot. Back in the days, a developer needed to care about context management, memory creation, and all of that. Right now, as I say, with harnesses such as Codex, all that is gone.
For Luis, that progress comes from combining more capable models with more powerful harnesses. Together, they allow engineers and developers to run increasingly complex workflows over much longer periods of time.

Agents should follow the same access rules as employees
When it comes to security, we wanted to understand the safest way to connect AI agents to company data without creating new permission and access risks.
Luis explained it through a familiar workplace example. When a company onboards someone who works in finance, it gives that person access to financial data, but not to HR data they do not need.
He argues that companies can apply the same principle to AI agents, using the same access controls and permission rules they already use for employees:
In the same way that when you onboard an employee into your company, and let’s say he or she is working in finance, you only give access to the finance data. You don’t give access to the HR data, right? So, in the same way, those same principles, those same ACLs, can also apply to AI agents, right?
For him, the principle comes down to giving agents access only to what they require to perform their tasks:
So, in my view, it’s about having the roles, the permissions, and the setup for the agents to only see what they need to see for completing their workflow, basically.
Evals are fundamental to AI development
Finally, we asked Luis what developers should test and measure before releasing an AI agent to real users.
Once again, he came back to evals. He said he keeps returning to the concept because it remains critical and fundamental to AI development:
I think we go back to evals once again, right? And I keep coming back to that concept because it’s absolutely critical and fundamental for AI development, right? If you’re able to put together that representative eval set containing easy cases, medium cases, and also edge cases, and you’re able to highly lean on your skills, on your guardrails, on the prompts you send to the model, that will give you the guarantees that when you deploy to production, nothing will break.
He added that evals also give developers a way to test for regressions when they upgrade to a new model and make sure the change does not break existing workflows:
Furthermore, when you upgrade your model to a new one, you have a way to test regressions and make sure that that upgrade won’t break anything.
System design taste will never go away
Before joining OpenAI, Luis also worked at Google, so over the years he has seen plenty of developers, roles, and career dilemmas up close. That is why, to close the interview, we asked him which skills developers should focus on today to stay relevant in the age of AI agents.
He believes that, in this new era, engineers will no longer create code as their primary output, because code has become very cheap to produce.
Instead, he sees much more value in everything developers build around the model:
I think now, in this new era we’re entering, the primary thing that engineers are creating is not code anymore. Code is very cheap to produce. It’s all about building the system around it, so the system of intent, the guardrails, the constraints for the agent to work reliably, right? So I think creating those harnesses around the model will be absolutely critical.
Luis ended our interview with a simple but sharp statement:
Plus, I always say, like, system design taste will never go away.


