<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>ShiftMag</title>
	<atom:link href="https://shiftmag.dev/feed/" rel="self" type="application/rss+xml" />
	<link>https://shiftmag.dev/</link>
	<description>Insightful engineering content &#38; community</description>
	<lastBuildDate>Tue, 06 Oct 2026 10:25:47 +0000</lastBuildDate>
	<language>en-GB</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://shiftmag.dev/wp-content/uploads/2026/09/cropped-Shift-32x32.webp</url>
	<title>ShiftMag</title>
	<link>https://shiftmag.dev/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>QA Is What Happens When Software Meets Reality</title>
		<link>https://shiftmag.dev/qa-is-what-happens-when-software-meets-reality-11744/</link>
		
		<dc:creator><![CDATA[Aldijana Culezovic]]></dc:creator>
		<pubDate>Tue, 06 Oct 2026 10:25:12 +0000</pubDate>
				<category><![CDATA[Quality Assurance]]></category>
		<category><![CDATA[QA]]></category>
		<category><![CDATA[software testing]]></category>
		<category><![CDATA[UX]]></category>
		<guid isPermaLink="false">https://shiftmag.dev/?p=11744</guid>

					<description><![CDATA[<p>Developers can prove the code works in a controlled environment. QA checks whether it survives the messy, unpredictable world customers actually live in.</p>
<p>The post <a href="https://shiftmag.dev/qa-is-what-happens-when-software-meets-reality-11744/">QA Is What Happens When Software Meets Reality</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></description>
										<content:encoded><![CDATA[<figure class="wp-block-post-featured-image"><img fetchpriority="high" decoding="async" width="1200" height="630" src="https://shiftmag.dev/wp-content/uploads/2026/09/developer-vs.-QA.png?x32039" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="" style="object-fit:cover;" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/developer-vs.-QA.png 1200w, https://shiftmag.dev/wp-content/uploads/2026/09/developer-vs.-QA-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/developer-vs.-QA-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/developer-vs.-QA-768x403.png 768w" sizes="(max-width: 1200px) 100vw, 1200px" /></figure>


<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">The developer already tested it, so why do you need to test it too?</p>
</blockquote>



<p class="wp-block-paragraph">If you work in tech, <strong>you’ve probably heard this sentence a thousand times</strong>.</p>



<p class="wp-block-paragraph">A developer writes the code, runs the checks, all the tests pass, and the feature works.</p>



<p class="wp-block-paragraph">So why do companies still hire Quality Assurance (QA) engineers? Are we just here to double-check developers’ work, or is there a deeper purpose?</p>



<p class="wp-block-paragraph">The short answer is: no.</p>



<p class="wp-block-paragraph">The longer answer is that <strong>there’s a fundamental difference between writing code and creating a great customer experience</strong>. The value of QA lies in perspective, risk awareness, and understanding how people actually use a product.</p>



<h2 class="wp-block-heading"><span id="the-builder">The builder</span></h2>



<p class="wp-block-paragraph">To build a truly great product, you need <strong>two very different mindsets</strong>: the builder and the chaos maker.</p>



<p class="wp-block-paragraph">A developer’s brain is wired for creation, logic, and structure. While QA starts the day with a finished product, a developer starts with a blank screen. Their job is to create something out of nothing and assemble the many moving parts of a complex system into something that works exactly as intended.</p>



<p class="wp-block-paragraph"><strong>Developers are responsible for creating the magic</strong>. To build even a single feature, they have to master a labyrinth of tools, frameworks, databases, and configurations. They carry the responsibility of making sure the architecture is secure, the system can handle heavy load, and the new code integrates cleanly with millions of lines of existing logic. Their work is highly complex and demands deep understanding, precision, and sustained focus. They are the architects of the digital world.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Because their primary mission is building and stabilizing the system, their testing mindset naturally reflects that. When developers test their code, they are checking whether their blueprint holds up. They focus on the happy path, along with a few common scenarios, in a controlled environment, using standard data to confirm that the core functionality works as intended. Their attention is on making sure the foundation is solid.</p>
</blockquote>



<p class="wp-block-paragraph">Think of a developer as an architect and engineer combined. They design a beautiful, highly complex door, check that the handle turns smoothly, inspect the hinges, and confirm that it opens perfectly.</p>



<p class="wp-block-paragraph"><strong>Now what? Aren’t developers already responsible for the whole thing?</strong></p>



<h2 class="wp-block-heading"><span id="the-chaos-maker">The chaos maker</span></h2>



<p class="wp-block-paragraph">While developers are incredibly skilled at what they do, they cannot be expected to carry the entire burden of the user experience alone. That is where QA steps in &#8211; to relieve some of that pressure and help refine the product for the real world.</p>



<p class="wp-block-paragraph">If a developer’s brain is wired for creation, a <strong>QA’s brain is wired for exploration and chaos</strong>. The goal is not to prove that the software works, but to discover what happens when the scenario is not perfect. We test unusual cases, unpredictable behavior, and the edge cases that real users may actually trigger. We test for the messy real world.</p>



<p class="wp-block-paragraph">Because of this, QA goes far beyond engineering. The work calls for empathy, human psychology, and an eye for the emotional, reliability, and invisible factors that make customers want to use a product.</p>



<p class="wp-block-paragraph">As a QA, <strong>I’m constantly stepping into the user’s shoes</strong> and asking myself a very different set of questions:</p>



<ul class="wp-block-list">
<li>Is this accessible and intuitive to use?</li>



<li>What emotion does this product create for me &#8211; frustration or delight?</li>



<li>Would I still enjoy using it tomorrow, or would it already be starting to annoy me?</li>
</ul>



<h2 class="wp-block-heading"><span id="what-happens-in-the-real-world">What happens in the real world</span></h2>



<p class="wp-block-paragraph">To make this point come to life, let me share a personal experience. </p>



<p class="wp-block-paragraph">I’ve seen that <strong>even a large team of skilled developers means every single scenario is not automatically covered</strong>. Tight deadlines, heavy workloads, and the sheer mental focus required to build complex architecture make it nearly impossible for any single perspective to catch everything. Once a feature is done with implementation, it’s QA’s turn to step in, look beyond the happy path, and hunt down every possible edge case.</p>



<h3 class="wp-block-heading">The &#8220;blank space&#8221; API trap</h3>



<p class="wp-block-paragraph">For instance, when new feature was implemented, it was time to test it. My team and I ran automated tests for our Voice APIs.</p>



<p class="wp-block-paragraph">During testing, <strong>I decided to try something slightly unusual</strong>: I sent a call using a sender name that contained a space. Immediately, the calls started failing. This turned out to be a critical issue. If a real client wanted to do something as simple and common as putting a space in their sender name (e.g., &#8220;Global Bank&#8221; instead of &#8220;GlobalBank&#8221;), their entire voice campaign would fail. We only discovered this flaw because we had an automated QA project specifically designed to stress-test the exact parameters our clients use daily, which even shows the importance of regression tests and their benefits before the release.</p>



<h3 class="wp-block-heading"><span id="the-edge-case-that-shouldn%e2%80%99t-break-the-system">The edge case that shouldn’t break the system</span></h3>



<p class="wp-block-paragraph">Another instance happened while I was testing the delivery reports for our system. The documentation stated that the record limit field was limited to 1 to 1000. My QA brain asked, &#8220;What happens if a user accidentally types 0?&#8221; <strong>When I entered 0, the system crashed and returned a 500 Internal Server Error</strong>. </p>



<p class="wp-block-paragraph">Technically, it was just a wrong number. But from a customer experience standpoint, this is a big deal. A 500 error tells the client, &#8220;Your servers are broken,&#8221; which causes panic, especially during real-time actions. It should have gracefully returned a 400 Bad Request, which politely tells the user, &#8220;Oops, your input was invalid, please try again.&#8221; </p>



<p class="wp-block-paragraph">Logically, it makes no sense for a user to request 0 records. But it is our duty as QA to make sure that even when a user does something illogical, the system handles it elegantly instead of catching on fire.</p>



<h3 class="wp-block-heading"><span id="the-trapped-user">The trapped user</span></h3>



<p class="wp-block-paragraph">The chaos isn’t always in the backend, either. Recently, while performing manual UI testing, I found a scenario where a user could fill out an entire complex form, only to find that the <strong>&#8220;Save&#8221; button remained permanently disable</strong>d because of a tiny edge-case conflict in the UI logic. Imagine the frustration of spending ten minutes typing out intricate details, only to be trapped at the finish line and unable to save the work. </p>



<p class="wp-block-paragraph">All of these instances prove why additional testing is necessary. It is incredibly easy for crucial edge cases to slip through when a team’s primary focus is building complex core functionality. No one should set a limit to 0. A sender name shouldn’t break an API. But humans aren’t strictly logical, and neither is the real world. That is why QA exists.</p>



<h2 class="wp-block-heading"><span id="qa-and-devs-are-two-halves-of-the-same-brain">QA and devs are two halves of the same brain</span></h2>



<p class="wp-block-paragraph">Ultimately, this story comes back to the same point: success. Despite our different approaches, <strong>both teams are chasing the same finish line</strong>. We come together, we collaborate, and we bring an idea to life in the best possible way.</p>



<p class="wp-block-paragraph">Both developers and QA want the same thing: to deliver a product that is simple, easy to use, unique, and highly functional. We want to build something so reliable and intuitive that the end user does not even have to think about how it works under the hood &#8211; it just does.</p>



<p class="wp-block-paragraph">Because my job involves finding edge cases, logging bugs, and occasionally breaking a developer’s hard work, it is easy for outsiders to assume that developers and QA are at war. But that is far from the truth. <strong>We are not enemies; we are an essential partnership</strong>. We are a continuous feedback loop where logic meets empathy.</p>



<p class="wp-block-paragraph">Just like a human being cannot function with only half a brain, a truly great product cannot survive in the wild with only one mindset. It needs the builder’s solid foundation, and it needs the chaos maker’s reality check.</p>



<p class="wp-block-paragraph"><strong>The developer brain ensures the software is built right. The QA brain ensures we’re building the right software for the customer.</strong> Together we create the product that customers love.</p>



<h2 class="wp-block-heading">Before you go&#8230;</h2>



<p class="wp-block-paragraph">Next time someone asks, &#8220;The developer already tested it, why do you need to test it too?&#8221; you’ll know exactly what to say: </p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Code lives on servers, but customer experience lives in the real world, and QA is the bridge between the two.</p>
</blockquote>
<p>The post <a href="https://shiftmag.dev/qa-is-what-happens-when-software-meets-reality-11744/">QA Is What Happens When Software Meets Reality</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>OpenAI&#8217;s Luis Velasco: Code is very cheap to produce, but the system around it isn&#8217;t</title>
		<link>https://shiftmag.dev/openai-fde-code-is-very-cheap-to-produce-12341/</link>
		
		<dc:creator><![CDATA[Anastasija Uspenski]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 13:20:51 +0000</pubDate>
				<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Event]]></category>
		<category><![CDATA[Infobip Shift 2026]]></category>
		<category><![CDATA[Luis Velasco]]></category>
		<category><![CDATA[OpenAI]]></category>
		<category><![CDATA[Video]]></category>
		<guid isPermaLink="false">https://shiftmag.dev/?p=12341</guid>

					<description><![CDATA[<p>Luis Velasco works as a Forward Deployed Engineer at OpenAI. For him, code is becoming very cheap to produce, while building the system around the model is becoming critical.</p>
<p>The post <a href="https://shiftmag.dev/openai-fde-code-is-very-cheap-to-produce-12341/">OpenAI&#8217;s Luis Velasco: Code is very cheap to produce, but the system around it isn&#8217;t</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">At the <a href="https://shiftmag.dev/tag/infobip-shift-2026/" target="_blank" rel="noreferrer noopener">Shift conference in Zadar</a> this September, we sat down with <a href="https://shiftmag.dev/openai-says-the-real-work-in-ai-coding-is-building-the-system-around-it-11976/" target="_blank" rel="noreferrer noopener">Luis Velasco, a Forward Deployed Engineer at OpenAI</a>, who describes his role simply:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">We help companies put AI to work in real environments.</p>
</blockquote>



<p class="wp-block-paragraph">We spoke with him about what it takes to move AI agents into production, how developers can make them reliable and secure, and which skills will matter most as AI changes software development.</p>



<h3 class="wp-block-heading"><span id="production-starts-with-evals"><strong>Production starts with evals</strong></span></h3>



<p class="wp-block-paragraph">As Luis hinted in the intro, Forward Deploy Engineers (FDE) help companies turn AI demos into real products. So we asked him what developers most often underestimate when they make that jump.</p>


<figure class="wp-block-post-featured-image"><img decoding="async" width="1200" height="630" src="https://shiftmag.dev/wp-content/uploads/2026/09/luis.png?x32039" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="" style="object-fit:cover;" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/luis.png 1200w, https://shiftmag.dev/wp-content/uploads/2026/09/luis-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/luis-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/luis-768x403.png 768w" sizes="(max-width: 1200px) 100vw, 1200px" /></figure>


<p class="wp-block-paragraph">According to him, <strong>the real challenge starts when a demo reaches production</strong>. A demo can work well in controlled conditions, but production forces teams to understand exactly what works, what does not, and why failures happen.</p>



<p class="wp-block-paragraph">That is where eval-driven (Evaluation-driven) development becomes important. He says OpenAI uses it extensively in its engagements to track and measure different workflows and tasks:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">We have a scientific approach to measure the performance of the models doing a task. When, basically, you clear those benchmarks, those evals, you are good to go to production in a really good and safe way.</p>
</blockquote>



<h3 class="wp-block-heading"><span id="the-deployment-gap-is-the-real-challenge"><strong>The deployment gap is the real challenge</strong></span></h3>



<p class="wp-block-paragraph">We wanted to understand what makes an AI agent reliable enough for production, rather than something that only works well in a controlled environment.</p>



<p class="wp-block-paragraph">Luis said<strong> it all comes down to “cracking the evals.”</strong> As he explained, today’s models already have impressive intelligence, but putting them to work in a company’s own environment requires giving them access to the company’s data, procedures, and tools:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">You know, those three things, we kind of call the deployment gap, right? And this is what we&#8217;re trying to close as FD, basically.</p>
</blockquote>



<h3 class="wp-block-heading"><span id="agents-went-from-six-minute-to-16-hour-tasks"><strong>Agents went from six-minute to 16-hour</strong> tasks</span></h3>



<p class="wp-block-paragraph">But what happens when an agent works on a task for hours rather than minutes? How should developers think about errors, memory, context, and human approval as those workflows become longer?</p>



<p class="wp-block-paragraph">Luis answered with an example from his own experience. While checking the latest data from METR’s long time horizon benchmark, he was surprised by how quickly agents had improved. In just over three years, they went from <strong>handling tasks that would take a human around six minutes to tasks that would take around 16 hours.</strong></p>



<p class="wp-block-paragraph">That represents a massive jump, he said:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">On one hand, you have the models getting better at long-term coherency and persistence, but on the other hand, you also have the harnesses improving a lot. Back in the days, a developer needed to care about context management, memory creation, and all of that. Right now, as I say, with harnesses such as Codex, all that is gone.</p>
</blockquote>



<p class="wp-block-paragraph">For Luis, that progress comes from combining more capable models with more powerful harnesses. Together, they allow engineers and developers to run increasingly complex workflows over much longer periods of time.</p>



<figure class="wp-block-image size-large"><img decoding="async" width="1024" height="683" src="https://shiftmag.dev/wp-content/uploads/2026/09/SHIFT_MAINSTAGE-19-1024x683.jpg?x32039" alt="" class="wp-image-12352" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/SHIFT_MAINSTAGE-19-1024x683.jpg 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/SHIFT_MAINSTAGE-19-300x200.jpg 300w, https://shiftmag.dev/wp-content/uploads/2026/09/SHIFT_MAINSTAGE-19-768x512.jpg 768w" sizes="(max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Credit: Filip Popović</figcaption></figure>



<h3 class="wp-block-heading"><span id="agents-should-follow-the-same-access-rules-as-employees"><strong>Agents should follow the same access rules as employees</strong></span></h3>



<p class="wp-block-paragraph">When it comes to security, we wanted to understand the safest way to connect AI agents to company data without creating new permission and access risks.</p>



<p class="wp-block-paragraph">Luis explained it through a familiar workplace example. When a company onboards someone who works in finance, it gives that person access to financial data, but not to HR data they do not need.</p>



<p class="wp-block-paragraph">He argues that companies can apply the same principle to AI agents, using the same access controls and permission rules they already use for employees:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">In the same way that when you onboard an employee into your company, and let&#8217;s say he or she is working in finance, you only give access to the finance data. You don&#8217;t give access to the HR data, right? So, in the same way, those same principles, those same ACLs, can also apply to AI agents, right?</p>
</blockquote>



<p class="wp-block-paragraph">For him, the principle comes down to giving agents access only to what they require to perform their tasks:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">So, in my view, it&#8217;s about having the roles, the permissions, and the setup for the agents to only see what they need to see for completing their workflow, basically.</p>
</blockquote>



<h3 class="wp-block-heading"><span id="evals-are-fundamental-to-ai-development"><strong>Evals are fundamental to AI development</strong></span></h3>



<p class="wp-block-paragraph">Finally, we asked Luis what developers should test and measure before releasing an AI agent to real users.</p>



<p class="wp-block-paragraph">Once again, he came back to evals. He said he keeps returning to the concept because it remains critical and fundamental to AI development:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">I think we go back to evals once again, right? And I keep coming back to that concept because it&#8217;s absolutely critical and fundamental for AI development, right? If you&#8217;re able to put together that representative eval set containing easy cases, medium cases, and also edge cases, and you&#8217;re able to highly lean on your skills, on your guardrails, on the prompts you send to the model, that will give you the guarantees that when you deploy to production, nothing will break.</p>
</blockquote>



<p class="wp-block-paragraph">He added that evals also give developers a way to test for regressions when they upgrade to a new model and make sure the change does not break existing workflows:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Furthermore, when you upgrade your model to a new one, you have a way to test regressions and make sure that that upgrade won&#8217;t break anything.</p>
</blockquote>



<h3 class="wp-block-heading"><span id="system-design-taste-will-never-go-away"><strong>System design taste will never go away</strong></span></h3>



<p class="wp-block-paragraph">Before joining OpenAI, Luis also worked at Google, so over the years he has seen plenty of developers, roles, and career dilemmas up close. That is why, to close the interview, we asked him which skills developers should focus on today to stay relevant in the age of AI agents.</p>



<p class="wp-block-paragraph">He believes that, in this new era, engineers will no longer create code as their primary output, because <strong>code has become very cheap to produce</strong>.</p>



<p class="wp-block-paragraph">Instead, he sees much more value in everything developers build around the model:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">I think now, in this new era we&#8217;re entering, the primary thing that engineers are creating is not code anymore. Code is very cheap to produce. It&#8217;s all about building the system around it, so the system of intent, the guardrails, the constraints for the agent to work reliably, right? So I think creating those harnesses around the model will be absolutely critical.</p>
</blockquote>



<p class="wp-block-paragraph">Luis ended our interview with a simple but sharp statement:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Plus, I always say, like, system design taste will never go away.</p>
</blockquote>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Open AI on Why AI Demos Break in Production" width="500" height="281" src="https://www.youtube.com/embed/wCiFGPgtQvw?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>
<p>The post <a href="https://shiftmag.dev/openai-fde-code-is-very-cheap-to-produce-12341/">OpenAI&#8217;s Luis Velasco: Code is very cheap to produce, but the system around it isn&#8217;t</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Review fatigue is real. Here&#8217;s what my team did about it</title>
		<link>https://shiftmag.dev/review-fatigue-12276/</link>
		
		<dc:creator><![CDATA[Ivona Skorjanc]]></dc:creator>
		<pubDate>Tue, 29 Sep 2026 13:53:26 +0000</pubDate>
				<category><![CDATA[Productivity]]></category>
		<category><![CDATA[ai-assisted development]]></category>
		<category><![CDATA[code review]]></category>
		<category><![CDATA[developer workflow]]></category>
		<category><![CDATA[engineering culture]]></category>
		<category><![CDATA[Pair programming]]></category>
		<guid isPermaLink="false">https://shiftmag.dev/?p=12276</guid>

					<description><![CDATA[<p>When the volume of code outpaces the cognitive budget of the people asked to read it, fatigue happens. </p>
<p>The post <a href="https://shiftmag.dev/review-fatigue-12276/">Review fatigue is real. Here&#8217;s what my team did about it</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></description>
										<content:encoded><![CDATA[<figure class="wp-block-post-featured-image"><img loading="lazy" decoding="async" width="1200" height="630" src="https://shiftmag.dev/wp-content/uploads/2026/09/ai-code-review-pr-pile-1200x630-1.png?x32039" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="" style="object-fit:cover;" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/ai-code-review-pr-pile-1200x630-1.png 1200w, https://shiftmag.dev/wp-content/uploads/2026/09/ai-code-review-pr-pile-1200x630-1-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/ai-code-review-pr-pile-1200x630-1-768x403.png 768w, https://shiftmag.dev/wp-content/uploads/2026/09/ai-code-review-pr-pile-1200x630-1-1024x538.png 1024w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></figure>


<p class="wp-block-paragraph">I have never particularly liked being the code reviewer. </p>



<p class="wp-block-paragraph">It feels like getting all the responsibility but zero fun of actually writing the code. Even before most of the code became AI generated, the paradox was familiar to many: the more lines of code a PR has, the less likely it is that the changes will get a very detailed review. <strong>Why do we expect standard code reviews will still work when AI generates much more code volume?</strong></p>



<p class="wp-block-paragraph">For a workplace where agents are doing the heavy lifting, benefits of code review through classic PRs are gone, and there are a few reasons why:</p>



<ol start="1" class="wp-block-list">
<li>Engineers do more <strong>prompting and decision</strong> making than coding. But most people aren&#8217;t reviewing the former.</li>



<li>The <strong>reviewer might be the first one to actually have to read all the code</strong>. Since the implementer gets familiar with the implementation through the prompting process, <strong>the reviewer is in a worse position</strong>. If they go through everything line by line, it probably takes them longer than it took the original developer. If the real time cost of review isn&#8217;t reflected in sprint capacity planning, reviewers are probably not given enough time to finish the task.</li>



<li><strong>AI code review agents are a logical next step.</strong> To some extent, they are already utilized by AI-transformed companies. However, just throwing an agent at a PR without any structural and conceptual changes around code reviews, such as new types of safeguards and a robust production setup with automated rollbacks, misses the point of code reviews entirely. The human is no longer in the loop.</li>



<li><strong>Context switching.</strong> It was always an unwanted companion of software engineers, but now it’s like a shadow we can’t escape. While our agents are thinking and combobulating, we do something else. We spin up a few agents for a different task, and then another for the next one. What happens to the PRs then? They pile up. The pressure to downsize that pile can be real, so the review quality becomes questionable.</li>
</ol>



<p class="wp-block-paragraph">Teams have adapted to AI quickly in many parts of the workflow, but code review is still lagging behind. Instead of following the same workflow as before, we should rethink what code review is and how we practice it. While this might not suit best for every organization, for me and my team, the most useful model right now is a <a href="https://shiftmag.dev/tag/pair-programming/" data-type="link" data-id="https://shiftmag.dev/tag/pair-programming/">pair-programming</a> spinoff: <strong>pair planning and pair validation</strong>.</p>



<h2 class="wp-block-heading"><span id="pair-planning"><strong>Pair planning</strong></span></h2>



<p class="wp-block-paragraph">Pair planning is the first step. It starts with prompting the agent. And it matters how you do it.</p>



<p class="wp-block-paragraph">The easy option is to drop the task description to an AI agent and let it do the work however it wants. Long-term, this is a good way to create a big pile of slop.</p>



<p class="wp-block-paragraph">A better way is to ask the agent to <strong>propose multiple options and document them</strong>, without any code changes. While the LLM is doing its magic, this is your time to take a moment and think about the task in front of you. Ideally, the thinking happens out loud because you&#8217;re pairing with another colleague.</p>



<p class="wp-block-paragraph">Why is the part when you&#8217;re thinking without any AI assistance important? Once you start reading the AI output, it&#8217;s highly likely your brain will lock in. Alternatives which the agent didn&#8217;t propose are more difficult to think of. It&#8217;s also possible that the task description wasn’t detailed enough and the agent didn’t take an important detail into account. Any caveats, edge cases, deployment scenarios, or other details that come to mind while you&#8217;re brainstorming about the possible task implementation details are good to note down before you read the agent output.</p>



<p class="wp-block-paragraph">The agent is done planning. Output is ready. <strong>Now what?</strong></p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="538" src="https://shiftmag.dev/wp-content/uploads/2026/09/ai-code-review-pair-planning-1200x630-2-1024x538.png?x32039" alt="" class="wp-image-12284" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/ai-code-review-pair-planning-1200x630-2-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/ai-code-review-pair-planning-1200x630-2-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/ai-code-review-pair-planning-1200x630-2-768x403.png 768w, https://shiftmag.dev/wp-content/uploads/2026/09/ai-code-review-pair-planning-1200x630-2.png 1200w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Credit: Google AI</figcaption></figure>



<p class="wp-block-paragraph">Both you and your pair should get to reading. Flag any issues you see and discuss them. Refine the plan. Squeeze the implementation details from the agent. Be aware of all components that will change and how. This is the juicy part of the task.</p>



<p class="wp-block-paragraph">With AI added to the traditional code review flow, we often plan the implementation twice. The first time is by the original engineer implementer. This person will make all the decisions individually. They will guide the implementation. Then the second time is by the reviewer, <a href="https://shiftmag.dev/it-may-take-longer-to-review-a-pr-than-it-takes-to-write-it-9927/" data-type="link" data-id="https://shiftmag.dev/it-may-take-longer-to-review-a-pr-than-it-takes-to-write-it-9927/">possibly in even more detail</a>. If the reviewer doesn’t agree with the way task was implemented and proposes a different approach, the author and agent will go back to the beginning. All tokens used on the wrong path were wasted. </p>



<p class="wp-block-paragraph">These kind of &#8220;disposable code&#8221; scenarios are happening more and more. <strong>When pairing with another colleague during the planning phase, it&#8217;s more likely you will choose the best implementation option and reach your goal in minimal time and tokens spent.</strong></p>



<h2 class="wp-block-heading"><span id="structural-prompt-committing"><strong>Structural prompt committing</strong></span></h2>



<p class="wp-block-paragraph">Most of the engineering work gets done in the planning phase. Decisions are made and propagated into code. From commits only, we can only see <em>what</em> was decided. <em>Why</em> it was decided and <em>what the alternatives were</em> is valuable information that tends to get lost.</p>



<p class="wp-block-paragraph">Prompting practices can be refined to avoid losing decision data. I practice structural prompt committing. In the codebase, each task is documented with my prompts, distilled versions of agent responses, and a few more useful files. The structure can look something like:</p>



<ul class="wp-block-list">
<li>task folder
<ul class="wp-block-list">
<li><strong>prompt.md</strong>
<ul class="wp-block-list">
<li>chronological log of your prompts and distilled info about agent responses, append only</li>
</ul>
</li>



<li><strong>context.md</strong>
<ul class="wp-block-list">
<li>current context, changes can be traced through git versioning</li>



<li>either for human reference or for loading agent context</li>
</ul>
</li>



<li><strong>plan.md</strong>
<ul class="wp-block-list">
<li>after all the decisions are made via prompting, full plan is documented to a separate file</li>
</ul>
</li>



<li><strong>summary.md</strong>
<ul class="wp-block-list">
<li>summary of implementation changes</li>
</ul>
</li>
</ul>
</li>
</ul>



<p class="wp-block-paragraph">This structure is useful while working on the task, especially if multiple persons are involved. Afterwards it can be deleted. For important decisions, architectural decision records should be made.</p>



<h2 class="wp-block-heading"><span id="implementation-ideas"><strong>Implementation ideas</strong></span></h2>



<p class="wp-block-paragraph">How you decide to do the implementation after the plan is ready depends on many factors. Some ideas can be applied generally:</p>



<ul class="wp-block-list">
<li>If unsure between options proposed by an agent, <strong>ask it to prototype them on multiple branches</strong>. See how they act in action and validate through tests, metrics, etc.</li>



<li>Small commits are desirable. It is easier to follow the changes and revert precisely if needed.</li>



<li>For bigger tasks, create checkpoints. Don&#8217;t let the agents create a massive change which will have to be fully reverted if it proves wrong.</li>
</ul>



<p class="wp-block-paragraph">You might be wondering, <strong>what should you do while Claude is claudeing</strong>? Should the pair of colleagues just&#8230; wait and watch?</p>



<p class="wp-block-paragraph">Well, maybe. If the checkpoints are small enough it might make sense to wait and discuss the direction the agent is taking. If not, it could make sense to split up and meet again when the project is ready for validation and review. It&#8217;s up to you to weigh if the context switch is worth it.</p>



<h2 class="wp-block-heading"><span id="pair-validation"><strong>Pair validation</strong></span></h2>



<p class="wp-block-paragraph">Whatever road you choose, you should arrive at the last part: pair validation. Why didn’t I call it pair review? Because review sounds like you&#8217;re just reading the code. Instead, at this stage you should validate, i.e. prove, that code works.</p>



<p class="wp-block-paragraph">If the implementer and reviewer validate alone, the job is doubled and less efficient. <strong>Since AI is generating most code, we are all validators more than implementers. </strong>Let&#8217;s not validate twice. Two colleagues should go through the code changes in a structured way, making sure the requirements and acceptance criteria are met. Ideally, start from the tests. Are they proving the new code is working? Are they covering all use cases? Are they guarding the future implementation from possible breaking changes?</p>



<p class="wp-block-paragraph">Through conversation, concerns can be raised and resolved immediately. Learning and knowledge sharing are instantaneous. Don&#8217;t lose this value by spending time navigating in heaps of generated code and writing async comments someone should check between three different context-switching sessions.</p>



<p class="wp-block-paragraph">If you work in an environment that allows it, I encourage you to start thinking about the process of implementing and reviewing as one unit, not two separate jobs for two separate people. A good way to avoid getting in the “just review this quickly please” trap is to plan ahead. During planning of what each teammate will do in a given period, allocate the same amount of time and resources for both the implementer and reviewer. <strong>Teamwork begins before the task is started.</strong></p>



<p class="wp-block-paragraph"><strong>Not all jobs can follow this framework.</strong> Examples are open-source project PRs, distributed teams across multiple time zones, and contractor engagements. In these kinds of environments, reimagining how we do reviews gets even more interesting. I expect new and completely different frameworks to emerge. Until then, my team and I are quite confident shipping pair-planned and pair-validated code.</p>



<p class="wp-block-paragraph">Until then, my team and I are quite confident shipping pair-planned and pair-validated code.</p>
<p>The post <a href="https://shiftmag.dev/review-fatigue-12276/">Review fatigue is real. Here&#8217;s what my team did about it</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Netlify CTO: &#8220;No longer should my developers be writing code&#8221;</title>
		<link>https://shiftmag.dev/netlify-cto-no-longer-should-my-developers-be-writing-code-12298/</link>
		
		<dc:creator><![CDATA[Anastasija Uspenski]]></dc:creator>
		<pubDate>Mon, 28 Sep 2026 06:42:17 +0000</pubDate>
				<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Event]]></category>
		<category><![CDATA[Dana Lawson]]></category>
		<category><![CDATA[Infobip Shift 2026]]></category>
		<category><![CDATA[Netlify]]></category>
		<category><![CDATA[Video]]></category>
		<guid isPermaLink="false">https://shiftmag.dev/?p=12298</guid>

					<description><![CDATA[<p>Dana Lawson leads R&#038;D at Netlify, overseeing engineering, product, and design. From where she sits, she sees the real shift clearly: one person or a tiny team can now cover far more ground.</p>
<p>The post <a href="https://shiftmag.dev/netlify-cto-no-longer-should-my-developers-be-writing-code-12298/">Netlify CTO: &#8220;No longer should my developers be writing code&#8221;</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">We spoke with Dana at the <a href="https://shiftmag.dev/tag/infobip-shift-2026/" target="_blank" rel="noopener">Shift conference in Zadar</a>, where we discussed AI agents, the evolving role of developers, and the future of software teams.&nbsp;</p>



<h2 class="wp-block-heading"><span id="small-teams-are-taking-over-the-whole-build"><strong>Small teams are taking over the whole build</strong>&nbsp;</span></h2>



<p class="wp-block-paragraph">The structure of software teams is one area where Dana already sees a noticeable shift. Larger teams traditionally split these roles among several people, but much smaller groups can now handle them together:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">The biggest difference is the empowerment that teams have. Before, you would have teams of 5, 6, 8, 10 people that all had very distinct roles in the software development life cycle, and I think the biggest shift is the enablement that people that are enabled solo, or two or three, can do everything end-to-end.</p>
</blockquote>



<p class="wp-block-paragraph">That shift has also influenced how Netlify organizes product development. Dana says the company no longer relies on sequential handoffs between product managers, designers, and engineers.</p>



<p class="wp-block-paragraph">Instead, Netlify gives teams more autonomy through a narrative-driven product roadmap. The company sets the broader direction, while individual teams decide how to shape what they build next.&nbsp;</p>



<p class="wp-block-paragraph">To stay aligned without introducing additional layers of coordination, Netlify combines familiar tools and routines with AI agents:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">We really let our teams have some personal take on what they&#8217;re going to build next, but we utilize systems like Linear and stand-ups, and honestly agents to keep us all in line so that we can move at the speed of machines.&nbsp;</p>
</blockquote>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="538" src="https://shiftmag.dev/wp-content/uploads/2026/09/Dana1-1024x538.png?x32039" alt="" class="wp-image-12326" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/Dana1-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/Dana1-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/Dana1-768x403.png 768w, https://shiftmag.dev/wp-content/uploads/2026/09/Dana1.png 1200w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Credit: Filip Popović</figcaption></figure>



<h2 class="wp-block-heading"><span id="low-ego-leadership"><strong>Low-ego leadership</strong>&nbsp;</span></h2>



<p class="wp-block-paragraph">More autonomous teams also change what Netlify CTO expects from the people leading them.&nbsp;</p>



<p class="wp-block-paragraph">When asked which leadership lesson has stayed with her through roles at GitHub, Heptio, New Relic, and Netlify, Dana returned to two ideas, trust people and recognize expertise regardless of title:&nbsp;</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Low ego. You are going to have people smarter than you in the room just because you have the amazing title and you are marked the leader. Your job is to be humble, build connections, and really learn, and be a part of your team and not just mining your team, and I think the biggest lesson is to trust people and allow them to be adults.</p>
</blockquote>



<p class="wp-block-paragraph">That principle becomes even more relevant as the work itself changes and developers increasingly operate alongside AI agents.&nbsp;</p>



<h2 class="wp-block-heading"><span id="developer-experience-is-now-an-agent-experience"><strong>Developer experience is now an agent experience</strong>&nbsp;</span></h2>



<p class="wp-block-paragraph">When we asked Dana what a great developer experience looks like today, she framed the question around a different relationship between engineers and code.&nbsp;</p>



<p class="wp-block-paragraph">For her, the developer’s job is increasingly about enabling agents to act safely on their behalf:&nbsp;</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">A great developer experience is an agent experience. No longer should my developers be writing code. They should be enabling agents to do stuff safely and on their behalf.&nbsp;</p>
</blockquote>



<p class="wp-block-paragraph">That kind of environment also depends on developers being able to understand what those agents are doing.&nbsp;</p>



<p class="wp-block-paragraph">Dana points specifically to clarity, signals, and telemetry as the mechanisms engineers need both to help agents complete their work and to understand what they have done once changes reach production:&nbsp;</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">A good engineering experience is ensuring that we have clarity, signals, and telemetry so that their agents can get the job done, and so that they can know what the agents did in production.&nbsp;</p>
</blockquote>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="538" src="https://shiftmag.dev/wp-content/uploads/2026/09/Dana2-1024x538.png?x32039" alt="" class="wp-image-12331" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/Dana2-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/Dana2-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/Dana2-768x403.png 768w, https://shiftmag.dev/wp-content/uploads/2026/09/Dana2.png 1200w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Credit: Filip Popović</figcaption></figure>



<h2 class="wp-block-heading"><span id="from-writing-code-to-system-design"><strong>From writing code to system design</strong>&nbsp;</span></h2>



<p class="wp-block-paragraph">If agents take on more of the direct implementation work, Dana sees engineers spending more of their attention at the system level.&nbsp;<br>&nbsp;<br>She describes that change as a move away from creating code itself and toward designing the environment in which software is built and operated:&nbsp;</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">We&#8217;ve moved it from creating code to really architecting systems and allowing engineers to do what they do best, which is build trust within the systems.</p>
</blockquote>



<p class="wp-block-paragraph">In that framing, the impact of AI is not limited to making individual development tasks faster. It also changes where engineering judgment is applied.&nbsp;</p>



<h2 class="wp-block-heading"><span id="don%e2%80%99t-use-ai-agents-to-automate-bad-processes"><strong>Don’t use AI agents to automate bad processes</strong>&nbsp;</span></h2>



<p class="wp-block-paragraph">Dana is also cautious about one of the most straightforward ways companies can approach AI adoption,<strong> taking an existing process and simply inserting an agent into it</strong>.&nbsp;</p>



<p class="wp-block-paragraph">Instead, she argues that teams should reconsider the workflow itself:&nbsp;</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Agents can&#8217;t do everything. They can do everything, but they shouldn&#8217;t do everything, and it&#8217;s a trap. Don&#8217;t start building things on existing workflows. Change your workflows.</p>
</blockquote>



<p class="wp-block-paragraph">That becomes even more relevant as Dana expects engineering roles, tools, and workflows to change significantly over the next few years:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Enable an agent experience, because in two years, our jobs will significantly change, and some of the tools that we&#8217;ve been building and some of the workflows will be commoditized.&nbsp;</p>
</blockquote>



<h2 class="wp-block-heading"><span id="that-doesn%e2%80%99t-mean-replacing-every-existing-tool"><strong>That doesn’t mean replacing every existing tool</strong>&nbsp;</span></h2>



<p class="wp-block-paragraph">At the same time, Dana is not arguing that teams should replace technology simply because something newer exists.&nbsp;</p>



<p class="wp-block-paragraph">She specifically points out that tools such as Bash can still have a place when they remain appropriate for the problem.&nbsp;</p>



<p class="wp-block-paragraph"><strong>Her advice is:&nbsp;“Keep it simple.”&nbsp;</strong></p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" title="Netlify CTO: Developer Experience Is Now Agent Experience!" width="500" height="281" src="https://www.youtube.com/embed/ivj64wMJc2E?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
</div></figure>


<figure class="wp-block-post-featured-image"><img loading="lazy" decoding="async" width="1200" height="630" src="https://shiftmag.dev/wp-content/uploads/2026/09/dana.naslovna.png?x32039" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="" style="object-fit:cover;" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/dana.naslovna.png 1200w, https://shiftmag.dev/wp-content/uploads/2026/09/dana.naslovna-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/dana.naslovna-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/dana.naslovna-768x403.png 768w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></figure><p>The post <a href="https://shiftmag.dev/netlify-cto-no-longer-should-my-developers-be-writing-code-12298/">Netlify CTO: &#8220;No longer should my developers be writing code&#8221;</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>MongoDB: One write, two logs</title>
		<link>https://shiftmag.dev/mongodb-one-write-two-logs-12267/</link>
		
		<dc:creator><![CDATA[Frane Jelavic]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 09:04:30 +0000</pubDate>
				<category><![CDATA[Data]]></category>
		<category><![CDATA[database internals]]></category>
		<category><![CDATA[mongodb]]></category>
		<category><![CDATA[replication]]></category>
		<category><![CDATA[storage engines]]></category>
		<category><![CDATA[wiredtiger]]></category>
		<guid isPermaLink="false">https://shiftmag.dev/?p=12267</guid>

					<description><![CDATA[<p>At 55,000 operations per second across a 20 TB sharded cluster, you stop theorizing about databases and start reading their source code.</p>
<p>The post <a href="https://shiftmag.dev/mongodb-one-write-two-logs-12267/">MongoDB: One write, two logs</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></description>
										<content:encoded><![CDATA[<figure class="wp-block-post-featured-image"><img loading="lazy" decoding="async" width="1200" height="620" src="https://shiftmag.dev/wp-content/uploads/2026/09/mongodb-one-write-two-logs.png?x32039" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="" style="object-fit:cover;" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/mongodb-one-write-two-logs.png 1200w, https://shiftmag.dev/wp-content/uploads/2026/09/mongodb-one-write-two-logs-300x155.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/mongodb-one-write-two-logs-766x396.png 766w, https://shiftmag.dev/wp-content/uploads/2026/09/mongodb-one-write-two-logs-1024x529.png 1024w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></figure>


<p class="wp-block-paragraph">When I first encountered MongoDB, I tried to understand it through the database I knew best: PostgreSQL. My journey was a <em>baptism by fire</em>. I was managing a sharded cluster with 20 TB of data and as many as ~55k operations per second, including ~29k writes. The cluster had six shards, each with its own replica set, a config server replica set, and six routers.</p>



<p class="wp-block-paragraph">To operate that cluster safely, I had to stop treating MongoDB like PostgreSQL and learn how it worked internally.</p>



<p class="wp-block-paragraph">At one point, while tackling a flow control problem on the stated cluster, an AI model recommended increasing the portion of memory assigned to the WiredTiger cache. At the moment, the recommendation sounded genuine and correct, but I did not know what the model had based its conclusions on. To be completely honest, I didn’t even know how RAM is used by MongoDB. The nail in the coffin was the <a href="https://www.mongodb.com/docs/manual/core/wiredtiger/#memory-use">official docs</a> saying don’t change WiredTiger cache to RAM ratio.</p>



<p class="wp-block-paragraph">That experience clarified why I wanted to learn the internals. I want/need enough understanding to connect operational advice to a mechanism, identify the evidence it depends on, and decide. This applies whether the advice comes from AI, a colleague, documentation, or my own assumptions.</p>



<p class="wp-block-paragraph">What happens after a MongoDB client sends a write?</p>



<p class="wp-block-paragraph">How does that write pass through MongoDB and WiredTiger, and when does it become durable on disk?</p>



<p class="wp-block-paragraph">And when someone recommends changing a WiredTiger setting, what would have to be true for that recommendation to make sense?</p>



<h2 class="wp-block-heading"><span id="following-one-write">Following one write</span></h2>



<p class="wp-block-paragraph">When a client sends a write operation to a sharded cluster, the request first reaches a router called <code>mongos</code>. Routers cache data from the config server about chunk &#8211; shard placements. Using this data <code>mongos</code> determine which shard owns the relevant data and forward the operation to that shard&#8217;s primary node.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="972" src="https://shiftmag.dev/wp-content/uploads/2026/09/shardedCluster-1024x972.png?x32039" alt="" class="wp-image-12271" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/shardedCluster-1024x972.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/shardedCluster-300x285.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/shardedCluster-767x728.png 767w, https://shiftmag.dev/wp-content/uploads/2026/09/shardedCluster.png 1200w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Sharded MongoDB cluster&nbsp;</figcaption></figure>



<p class="wp-block-paragraph">However, this post does not compare replica sets and sharded clusters. Once routing is complete, the same question applies inside every shard.</p>



<p class="wp-block-paragraph">How does the primary <code>mongod</code> process request, persist it, and replicate it to secondary nodes?</p>



<h2 class="wp-block-heading"><span id="separation-of-concern">Separation of concern</span></h2>



<p class="wp-block-paragraph">To answer my question(s), I firstly had to understand where <code>mongod</code> ends and the storage engine <code>WiredTiger</code> begins.</p>



<p class="wp-block-paragraph"><code>mongod</code> handles client commands, authorization, query execution, transaction coordination, and replication. WiredTiger is the embedded storage engine that stores collection and index data. It provides local transactions, MVCC, caching, compression, journaling, checkpoints, and crash recovery.</p>



<pre class="wp-block-code"><code>MongoDB client
      │ write operation
      ▼
mongos query router
      │ routes to the owning shard
      ▼
shard primary (mongod)
      ├── authorization and command handling
      ├── query parsing and execution
      ├── transaction and session coordination
      ├── replication and oplog coordination
      └── storage-engine API
                  │
                  ▼
             WiredTiger
                  ├── local transactions and MVCC
                  ├── internal cache
                  ├── journal
                  ├── checkpoints
                  └── collection and index data in dbPath</code></pre>



<p class="wp-block-paragraph">Ok, but how does this come into play? Boundaries are not absolute.</p>



<p class="wp-block-paragraph">The oplog belongs to MongoDB’s (logical) replication model, but WiredTiger stores it alongside other collection data.</p>



<h2 class="wp-block-heading"><span id="one-write-two-logs">One write, two logs</span></h2>



<p class="wp-block-paragraph">At this point, I had another question. If MongoDB already has a <code>journal</code>, why does it also need an oplog?</p>



<p class="wp-block-paragraph">The WiredTiger journal is the closest equivalent to PostgreSQL&#8217;s WAL (Write Ahead Log). It records the changes needed to recover <strong>one</strong> MongoDB member (<em>keep this thought</em>) after a crash. Unlike PostgreSQL&#8217;s WAL, MongoDB does not use the journal as its replication stream.</p>



<p class="wp-block-paragraph">The <code>oplog</code> has a different purpose. It is a capped collection stored in the <code>local</code> database as <code>local.oplog.rs</code>. Secondary members copy its entries and apply them to their own data.</p>



<p class="wp-block-paragraph">This raises an important consistency question. What prevents MongoDB from committing a document change without its corresponding oplog entry?</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="597" src="https://shiftmag.dev/wp-content/uploads/2026/09/twoWritePaths-v2-1024x597.png?x32039" alt="" class="wp-image-12273" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/twoWritePaths-v2-1024x597.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/twoWritePaths-v2-300x175.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/twoWritePaths-v2-768x448.png 768w, https://shiftmag.dev/wp-content/uploads/2026/09/twoWritePaths-v2.png 1200w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Client write path&nbsp;</figcaption></figure>



<h3 class="wp-block-heading"><span id="data-folder">Data folder</span></h3>



<p class="wp-block-paragraph">Let’s examine what happens on one node and then additionally complicate the story by adding secondaries.</p>



<p class="wp-block-paragraph">To understand local durability, I first looked inside the <code>dbPath</code> of a MongoDB member. A simplified listing:</p>



<pre class="wp-block-code"><code>$ ls /data/mongodb/
collection-&lt;ident>.wt
index-&lt;ident>.wt
journal/
WiredTiger
WiredTigerHS.wt
WiredTiger.lock
WiredTiger.turtle
WiredTiger.wt</code></pre>



<p class="wp-block-paragraph">The exact files and their names depend on the MongoDB version and configuration, but these are the main WiredTiger files relevant to this discussion.</p>



<p class="wp-block-paragraph">The <code>collection-&lt;ident&gt;.wt</code> and <code>index-&lt;ident&gt;.wt</code> files are WiredTiger tables that store collection records and index entries.</p>



<p class="wp-block-paragraph">MongoDB’s durable catalog maps logical objects, such as a collection namespace, to these storage-engine identifiers. WiredTiger maintains its own metadata in <code>WiredTiger.wt</code>, including its tables, their configuration, and their latest checkpoints.</p>



<p class="wp-block-paragraph">The small <code>WiredTiger</code> text file identifies the WiredTiger version used to create the database:</p>



<pre class="wp-block-code"><code>WiredTiger
WiredTiger 10.0.2: (November 30, 2021)</code></pre>



<p class="wp-block-paragraph"><code>WiredTigerHS.wt</code> is the history store. It keeps older committed versions of records needed by MVCC readers, snapshots, etc. WiredTiger removes versions once no active reader or snapshot can require them.</p>



<p class="wp-block-paragraph"><code>WiredTiger.turtle</code> contains enough metadata to locate and open the latest checkpoint of <code>WiredTiger.wt</code>. It bootstraps the metadata table during startup and recovery.</p>



<pre class="wp-block-code"><code>WiredTiger version string
WiredTiger 10.0.2: (November 30, 2021)
WiredTiger version
major=10,minor=0,patch=2
file:WiredTiger.wt
...
checkpoint_lsn=(2248,16291200)
...</code></pre>



<p class="wp-block-paragraph">The <code>checkpoint_lsn</code> identifies a position in the WiredTiger journal using a log file number and byte offset. During recovery, WiredTiger opens the last consistent checkpoint and replays later journal records.</p>



<p class="wp-block-paragraph">The <code>journal/</code> directory contains those write-ahead log files. They protect changes made after the last checkpoint from being lost during an unexpected shutdown.</p>



<p class="wp-block-paragraph">Finally, <code>WiredTiger.lock</code> is the file on which WiredTiger acquires a lock to prevent two processes from opening the same database directory simultaneously. The lock state matters, not merely the existence of the file.</p>



<h3 class="wp-block-heading"><span id="the-wiredtiger-cache">The WiredTiger cache</span></h3>



<p class="wp-block-paragraph">WiredTiger does not modify the <code>collection-*.wt</code> and <code>index-*.wt</code> files directly for every client operation. It first reads the required pages into its internal cache and applies the change there as part of a storage transaction. Modified pages become dirty until WiredTiger reconciles and writes them to disk.</p>



<p class="wp-block-paragraph">This internal cache keeps uncompressed collection data and is separate from the operating system’s filesystem cache where compressed data is stored.</p>



<p class="wp-block-paragraph">MongoDB therefore benefits from <strong>both caches</strong>, WiredTiger and the OS cache.</p>



<p class="wp-block-paragraph">Giving all available memory to the WiredTiger cache would leave too little memory for the filesystem cache and the rest of <code>mongod</code> and respecting the default, <code>max((availRAM - 1024) x 0.5, 256 MB)</code>, is strongly advised by <a href="https://www.mongodb.com/docs/manual/core/wiredtiger/#memory-use">the official docs.</a></p>



<p class="wp-block-paragraph">This is already enough to show why “increase the WiredTiger cache” is incomplete advice. A larger cache does not create memory. It transfers memory away from the filesystem cache and other allocations. Whether that trade is useful depends on the workload and the source of the observed pressure.</p>



<p class="wp-block-paragraph">When the cache approaches its limits, WiredTiger evicts pages to make room. Clean pages can be discarded and read again later. Dirty pages must first be reconciled into an on-disk representation.</p>



<p class="wp-block-paragraph">A committed change does not have to wait for its dirty page to reach the collection file. The journal supplies durability between checkpoints. This separation is the reason a write can be durable even though its final data page is still dirty in memory.</p>



<h3 class="wp-block-heading"><span id="recovering-after-a-crash">Recovering after a crash</span></h3>



<p class="wp-block-paragraph">In case of a crash, WiredTiger starts from the last complete checkpoint. The checkpoint metadata identifies a consistent view of the WiredTiger tables and the journal position associated with it.</p>



<p class="wp-block-paragraph">WiredTiger then replays the journal records created after that checkpoint. At a high level, recovery combines two inputs:</p>



<p class="wp-block-paragraph">&nbsp;last complete checkpoint + later durable journal records = recovered local state &nbsp;</p>



<h3 class="wp-block-heading"><span id="checkpoints-and-data-files">Checkpoints and data files</span></h3>



<p class="wp-block-paragraph">WiredTiger normally creates a checkpoint every (configurable default) 60 seconds. A checkpoint can take longer when there is more dirty data or the storage device is under pressure.</p>



<p class="wp-block-paragraph">During a checkpoint, WiredTiger writes a point-in-time view of dirty pages to the collection, index, history-store, and metadata files.</p>



<pre class="wp-block-code"><code>WiredTiger transaction
        │
        ├── dirty pages in cache ──checkpoint──&#x25b6; .wt data files
        │
        └── journal records ───────────────────&#x25b6; journal files

                         crash
                           │
                           ▼
       latest checkpoint + subsequent journal records
                           │
                           ▼
                    recovered state
</code></pre>



<h3 class="wp-block-heading"><span id="how-the-primary-records-an-operation">How the primary records an operation</span></h3>



<p class="wp-block-paragraph">For an incoming query, <code>mongod</code> will, among other things, determine collection and index changes and construct the oplog logical entry.</p>



<p class="wp-block-paragraph">It uses WiredTiger to start/use a transaction and either commits both parts or commits neither. This prevents a successful document change from existing on the primary without an operation (oplog entry) that secondaries can replicate.</p>



<p class="wp-block-paragraph">The oplog entry contains an operation type, namespace, timestamp, and the data required to reproduce the change. Its timestamp provides an ordering position in the replica-set history. The entry describes a logical insert, update, delete, command, or transaction operation rather than the low-level page modifications stored in the WiredTiger journal.</p>



<pre class="wp-block-code"><code>rs_a &#91;direct: primary] local> db.oplog.rs.findOne()
{
  ts: Timestamp(...),
  t: NumberLong(...),
  op: "u",
  ns: "foo.bar",
  o: { $v: 2, diff: { ... } },
  o2: { _id: 123 }
}</code></pre>



<p class="wp-block-paragraph">This is not a complete definition of how oplog and collection changes work, but the important invariant is: <strong>the primary commits the replicated data changes together with the oplog records which describe them.</strong></p>



<p class="wp-block-paragraph">Only the local commit is atomic. Sending or copying the oplog entry across the network is not part of that storage transaction.</p>



<h3 class="wp-block-heading"><span id="how-secondaries-copy-and-apply-oplog-entries">How secondaries copy and apply oplog entries</span></h3>



<p class="wp-block-paragraph">Each secondary continuously selects a sync source and streams newer oplog entries from it. The sync source is often the primary.</p>



<p class="wp-block-paragraph">The secondary writes fetched entries to its own oplog and then applies the represented operations to its local collections and indexes.</p>



<p class="wp-block-paragraph">The secondary uses its own WiredTiger instance to apply changes. They modify its cache, produce local journal records, and later become part of its checkpoints.</p>



<p class="wp-block-paragraph">This is the important distinction and the answer why we can’t just ship primaries <code>*.wt</code> binaries to the secondary.</p>



<p class="wp-block-paragraph">On two MongoDB servers, these are different files.</p>



<h3 class="wp-block-heading"><span id="oplog-versus-journal">Oplog versus journal</span></h3>



<figure class="wp-block-table"><table class="has-fixed-layout"><tbody><tr><td><strong>Property</strong></td><td><strong>Oplog</strong></td><td><strong>Journal</strong></td></tr><tr><td>Layer</td><td>MongoDB</td><td>WiredTiger</td></tr><tr><td>Purpose</td><td>Replication</td><td>Crash recovery</td></tr><tr><td>Contents</td><td>Logical MongoDB operations</td><td>Storage-engine recovery records</td></tr><tr><td>Consumed by</td><td>Secondary nodes</td><td>WiredTiger during startup/recovery</td></tr><tr><td>Scope</td><td>ReplicaSet logical history</td><td>One server only</td></tr><tr><td>Replicated?</td><td>Yes, logically</td><td>No</td></tr></tbody></table></figure>



<h3 class="wp-block-heading"><span id="takeaways">Takeaways</span></h3>



<p class="wp-block-paragraph">This gives a practical sequence for evaluating operational changes:</p>



<ol start="1" class="wp-block-list">
<li>Identify which component owns the behavior</li>



<li>Describe the mechanism that is a cause of the problem</li>



<li>Find evidence supporting that mechanism</li>



<li>Clarify what the improvements should be</li>



<li>Measure</li>



<li>Iterate</li>
</ol>



<p class="wp-block-paragraph">The technical facts in this article support that way of reasoning:</p>



<ul class="wp-block-list">
<li>The WiredTiger journal and the oplog solve different problems. The journal recovers one member; the oplog replicates logical operations between members.</li>



<li>MongoDB commits a replicated data change and its oplog entry in the same local storage transaction on the primary.</li>



<li>A durable change can still exist as a dirty cache page because the journal protects it between checkpoints.</li>



<li>Primaries and secondaries have separate WiredTiger databases, caches, journals, and data files.</li>



<li>Memory assigned to WiredTiger is part of a larger allocation decision that includes the filesystem cache and other process memory.</li>
</ul>



<h2 class="wp-block-heading"><span id="why-i-wanted-to-understand-this">Why I wanted to understand this</span></h2>



<p class="wp-block-paragraph">Learning how WiredTiger works does not mean that I can derive every production decision from first principles. It gives me a way to examine a recommendation instead of accepting it because the source sounds confident, and we all know AI can sound like that.</p>



<p class="wp-block-paragraph">The quality of an (AI) recommendation depends on the question and the evidence supplied to it.</p>



<h2 class="wp-block-heading"><span id="further-reading">Further reading</span></h2>



<ul class="wp-block-list">
<li><a href="https://www.mongodb.com/docs/manual/core/wiredtiger/">MongoDB manual: WiredTiger storage engine</a></li>



<li><a href="https://www.mongodb.com/docs/manual/core/journaling/">MongoDB manual: Journaling</a></li>



<li><a href="https://www.mongodb.com/docs/manual/core/replica-set-oplog/">MongoDB manual: Replica set oplog</a></li>



<li><a href="https://www.mongodb.com/docs/manual/core/replica-set-sync/">MongoDB manual: Replica set synchronization</a></li>



<li><a href="https://github.com/mongodb/mongo/blob/master/src/mongo/db/storage/README.md">MongoDB source documentation: Storage engine</a></li>
</ul>



<p class="wp-block-paragraph"><strong><em>A note on scope</em></strong> <em>&#8211; This article records my current understanding of MongoDB and WiredTiger, built through hands-on operation and continued study. I have tried to keep the technical details accurate, but any errors or oversimplifications are mine. <a href="https://www.mongodb.com/docs/manual/">MongoDB official docs</a> and <a href="https://source.wiredtiger.com/develop/index.html">WiredTiger documentation</a> should always be consulted for correctness and clarity.</em></p>
<p>The post <a href="https://shiftmag.dev/mongodb-one-write-two-logs-12267/">MongoDB: One write, two logs</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Runbooks Are Dead; Make Operations Executable</title>
		<link>https://shiftmag.dev/runbooks-are-dead-make-operations-executable-11880/</link>
		
		<dc:creator><![CDATA[Toni Babic]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 10:53:21 +0000</pubDate>
				<category><![CDATA[Software Engineering]]></category>
		<category><![CDATA[DevOps]]></category>
		<category><![CDATA[runbooks]]></category>
		<guid isPermaLink="false">https://shiftmag.dev/?p=11880</guid>

					<description><![CDATA[<p>Static runbooks break under pressure. Operational knowledge should live in executable tools instead of documents nobody can follow at 3 AM.</p>
<p>The post <a href="https://shiftmag.dev/runbooks-are-dead-make-operations-executable-11880/">Runbooks Are Dead; Make Operations Executable</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></description>
										<content:encoded><![CDATA[<figure class="wp-block-post-featured-image"><img loading="lazy" decoding="async" width="1200" height="630" src="https://shiftmag.dev/wp-content/uploads/2026/09/from-runbooks-to-tools-1.png?x32039" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="" style="object-fit:cover;" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/from-runbooks-to-tools-1.png 1200w, https://shiftmag.dev/wp-content/uploads/2026/09/from-runbooks-to-tools-1-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/from-runbooks-to-tools-1-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/from-runbooks-to-tools-1-768x403.png 768w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></figure>


<p class="wp-block-paragraph">The best runbook I ever read was also <strong>the most dangerous one I ever followed</strong>.</p>



<p class="wp-block-paragraph">It was about database failover procedure. 14 steps. Step 3 said &#8220;verify the replica is caught up&#8221; but didn&#8217;t say how. Step 7 said &#8220;if the promotion fails, escalate to the DBA team&#8221; like the DBA team was sitting around at 3 AM. Step 11 was just the word &#8220;validate&#8221; followed by a period.</p>



<p class="wp-block-paragraph">The whole document was a confession. It said: &#8220;We know this could go wrong, someone wrote down what they remembered, and we hope you figure out the rest&#8221;.</p>



<p class="wp-block-paragraph">This is the standard state of operational knowledge at most companies and <strong>in the AI era it&#8217;s an architectural dead end</strong>.</p>



<h2 class="wp-block-heading"><span id="the-runbook-is-the-institutional-debt">The runbook is the institutional debt</span></h2>



<p class="wp-block-paragraph"><strong>Every runbook starts with good intentions</strong>: something breaks, someone fixes it and they write down the steps so the next person doesn&#8217;t suffer. Six months later, half the commands are deprecated, the dashboard URLs point to a monitoring tool that was replaced, and the person who wrote it left the company.</p>



<p class="wp-block-paragraph">Think about what a runbook really asks of the person on call. At 3 AM, with adrenaline pumping and Slack pinging, you&#8217;re supposed to read a wall of text, mentally parse conditional branches (&#8220;if the primary is unreachable but the replica is healthy, proceed to step 4, otherwise skip to step 8&#8221;), manually copy-paste commands into a terminal, and hope you don&#8217;t typo a hostname.</p>



<p class="wp-block-paragraph">Let&#8217;s be real: nobody practices runbooks and nobody rehearses them. Your first time through the procedure is the incident itself. <strong>This looks like a runbook, but it behaves like a liability with good formatting</strong>.</p>



<h2 class="wp-block-heading"><span id="knowledge-should-be-executable"><strong>Knowledge should be executable</strong></span></h2>



<p class="wp-block-paragraph">Here’s the shift: stop asking where to document the procedure, and start asking how to build it into the platform.</p>



<p class="wp-block-paragraph">An executable runbook behaves like a service: you send it inputs, it runs the steps, and it returns a result. The real logic lives in tested, versioned code.</p>



<p class="wp-block-paragraph">The progression usually goes like this:</p>



<ul class="wp-block-list">
<li><strong>Stage 1: The prose runbook.</strong> A page someone wrote after the last outage. Accurate for about two weeks, then slowly rotting.</li>



<li><strong>Stage 2: The personal script.</strong> Someone on the team writes a bash script that automates the worst parts. It lives in their home directory on a bastion host. Nobody else knows it exists. When they leave, it&#8217;s gone.</li>



<li><strong>Stage 3: The shared tool.</strong> The script gets checked into a repo. It has a README. Maybe a Makefile. People use it, but they have to know where the repo is, clone it, install dependencies, and run it from the right machine.</li>



<li><strong>Stage 4: The config-driven service.</strong> The tool becomes a proper internal service with an API. It reads configuration from a central source of truth. It validates inputs before acting. It logs what it did. It can be called by a Slack bot, a CLI, or an AI agent without any of them knowing how it works internally.</li>



<li><strong>Stage 4 is where executable runbooks unlock the AI era</strong>. Here’s the hard part: AI agents can’t always follow a runbook. They can’t read a Grafana screenshot, SSH into a bastion host, or run a script from someone’s home directory. What they can do is call APIs, invoke tools, and pass structured inputs to services that return structured results.</li>
</ul>



<h2 class="wp-block-heading"><span id="during-incidents-the-system-should-do-the-procedure-for-you"><strong>During incidents, the system should do the procedure for you</strong></span></h2>



<p class="wp-block-paragraph">Let me make this concrete.</p>



<p class="wp-block-paragraph"><strong>Today</strong>. Today you get paged: the primary database is down. You open the internal wiki and start following the steps. Check replication lag. Copy the command. Swap in the hostname. Run it. Lag is zero. Good.</p>



<p class="wp-block-paragraph">Then you try to promote the replica, but it fails because you’re on the wrong jump host. So you find the right one, SSH in, run it again, and it works. You update the application connection string, forget the read-replica config, and suddenly half the services start failing because they’re still pointed at the old primary. You fix that too.</p>



<p class="wp-block-paragraph">Total time: 25 minutes. Steps you had to remember: 10. Things that went wrong because the runbook was out of date: 3.</p>



<p class="wp-block-paragraph"><strong>After</strong>. Same scenario. You get paged. You type <code>/db-failover</code> in Slack. The bot asks which cluster. You say <code>orders-db-prod</code>. It checks replication status, promotes the replica, updates the connection strings in the service mesh config, verifies all services can reach the new primary, and posts the result back to the incident channel.</p>



<p class="wp-block-paragraph">Total time: 90 seconds. Steps you had to remember: 1. Things that went wrong: 0.</p>



<p class="wp-block-paragraph">The executable version isn&#8217;t magic. Under the hood it&#8217;s calling the same commands. The difference is that the knowledge of <em>how</em> to do those steps, <em>in what order</em>, <em>with what validation between each one</em>, is encoded in a service rather than a document that someone reads for the first time during an outage.</p>



<h2 class="wp-block-heading"><span id="what-makes-it-a-tool-instead-of-a-script">What makes it a tool instead of a script</span></h2>



<p class="wp-block-paragraph">The real difference is whether the thing can be safely used at 3 AM by someone seeing it for the first time.</p>



<ul class="wp-block-list">
<li><strong>Every step validates before it acts</strong>. The tool checks replication health before touching anything. If active writes are still in flight, it aborts instead of corrupting data. The human doesn&#8217;t have to remember the preconditions. The tool refuses to proceed without them.</li>



<li><strong>Every step has a failure path</strong>. Promotion fails? Roll back. Connectivity check fails? Report which services are broken. The tool handles the branches. The person on call only needs to know one command.</li>



<li><strong>It&#8217;s idempotent</strong>.<strong> </strong>Running it twice on an already-promoted replica is a no-op, not a disaster. That matters at 3 AM when you&#8217;re not sure if the first attempt went through.</li>



<li><strong>One endpoint, every consumer. </strong><code>POST /tools/db-failover</code> is what the Slack bot calls, what the CLI calls, and what the MCP server exposes. One backend. One source of truth. Every surface gets the same capability.</li>
</ul>



<h2 class="wp-block-heading"><span id="why-this-matters-for-ai-agents">Why this matters for AI agents</span></h2>



<p class="wp-block-paragraph">This is the part I care about most. The current excitement around AI operations assistants, AI SRE copilots, AI incident responders, all of it hits the same wall: <strong>these systems can only operate on structured, machine-accessible interfaces</strong>.</p>



<p class="wp-block-paragraph">A document can be read and summarized, but it still gives an AI agent nothing to call.</p>



<p class="wp-block-paragraph">When you turn a runbook into a tool, you’re helping more than the person on call. You’re building something your AI agents can use too. Today it might run from a Slack command. Tomorrow it might be called by an MCP server or an agent. Either way, it just does the job.</p>



<p class="wp-block-paragraph">This is the bridge between the other articles in this series. <a href="https://shiftmag.dev/mcps-arent-apis-stop-treating-them-like-one-11420/" target="_blank" rel="noreferrer noopener">MCP servers need tools to expose</a>. <a href="https://shiftmag.dev/you-should-bring-your-ops-to-slack-11860/" target="_blank" rel="noreferrer noopener">Slack bots need tools to call </a>(article two). Executable runbooks are those tools. They&#8217;re the operational layer that both the human interface and the AI interface sit on top of.</p>



<h2 class="wp-block-heading"><span id="this-has-never-been-easier">This has never been easier</span></h2>



<p class="wp-block-paragraph">Five years ago, converting a runbook into an API-backed service meant writing a backend, wiring up auth, deploying it somewhere, and maintaining it forever. That&#8217;s why most teams never did it, the friction was too high.</p>



<p class="wp-block-paragraph">That equation has changed. Low-code and no-code automation tools like n8n have made it trivial to wrap existing APIs, databases, and CLI commands into callable workflows. You can build a db-failover tool in an afternoon by chaining HTTP requests, database queries, and conditional logic in a visual editor. Expose it as a webhook, and suddenly your Slack bot, your MCP server, and your AI agent all have something to call.</p>



<p class="wp-block-paragraph"><strong>The barrier isn&#8217;t technical anymore</strong>. We’re still defaulting to static documentation because that’s what we’ve always done. But when an n8n workflow can be built, tested, and deployed faster than you can format a runbook, “we don’t have time to automate it” stops being a valid excuse.</p>



<h2 class="wp-block-heading"><span id="start-with-the-procedure-that-scares-you-the-most">Start with the procedure that scares you the most</span></h2>



<p class="wp-block-paragraph">You don&#8217;t need to convert every runbook overnight. Start with the scariest runbook. Build the tool, test it in staging, use it in a maintenance window, then move on to the next one.</p>



<p class="wp-block-paragraph"><strong>Each tool you build reduces operational risk by a measurable amount</strong>. More importantly, each tool becomes a building block. Once db-failover exists as a tool, the full-site-recovery runbook doesn&#8217;t need to include database steps. It just calls db-failover and moves on.</p>



<p class="wp-block-paragraph">The end state is <strong>a platform where operational knowledge doesn&#8217;t live in documents</strong> <strong>but in services</strong>. Not because documents are bad, but because services are callable, testable, composable, and reachable by both humans and machines.</p>



<figure class="wp-block-table"><table class="has-fixed-layout"><tbody><tr><th class="has-text-align-left" data-align="left">Principle</th><th class="has-text-align-left" data-align="left">Why</th></tr><tr><td class="has-text-align-left" data-align="left">Prose runbooks are institutional debt</td><td class="has-text-align-left" data-align="left">They rot, require human interpretation under stress, and block AI integration entirely</td></tr><tr><td class="has-text-align-left" data-align="left">Knowledge should be executable</td><td class="has-text-align-left" data-align="left">Encode procedures in services that validate, log, and run consistently every time</td></tr><tr><td class="has-text-align-left" data-align="left">Expose tools through one API, not scripts scattered across repos</td><td class="has-text-align-left" data-align="left">A single backend serves Slack commands, CLI calls, and AI agent invocations</td></tr><tr><td class="has-text-align-left" data-align="left">Build for composition</td><td class="has-text-align-left" data-align="left">Each tool should be independently callable so higher-level runbooks can chain them</td></tr><tr><td class="has-text-align-left" data-align="left">Validate at every step, abort on failure</td><td class="has-text-align-left" data-align="left">The tool enforces preconditions. Humans under stress forget them</td></tr><tr><td class="has-text-align-left" data-align="left">Start with the scariest procedure</td><td class="has-text-align-left" data-align="left">Convert the runbook that causes the most anxiety first. Each tool reduces measurable risk</td></tr><tr><td class="has-text-align-left" data-align="left">Low-code tools eliminate the friction</td><td class="has-text-align-left" data-align="left">n8n and similar platforms let you build callable workflows in hours instead of weeks. &#8220;No time to automate&#8221; is no longer valid</td></tr></tbody></table></figure>



<p class="wp-block-paragraph">Humans still belong in operations; the machine should handle the repeatable work. Write the runbook once as code, and let the platform carry it.</p>
<p>The post <a href="https://shiftmag.dev/runbooks-are-dead-make-operations-executable-11880/">Runbooks Are Dead; Make Operations Executable</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>LLMs Have a New Limit &#8211; It Costs More to Think Longer</title>
		<link>https://shiftmag.dev/llms-have-a-new-limit-it-costs-more-to-think-longer-12083/</link>
		
		<dc:creator><![CDATA[Marko Crnjanski]]></dc:creator>
		<pubDate>Tue, 15 Sep 2026 12:50:23 +0000</pubDate>
				<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Event]]></category>
		<category><![CDATA[AI agents]]></category>
		<category><![CDATA[Infobip Shift 2026]]></category>
		<category><![CDATA[LLMs]]></category>
		<category><![CDATA[NVIDIA]]></category>
		<guid isPermaLink="false">https://shiftmag.dev/?p=12083</guid>

					<description><![CDATA[<p>NVIDIA’s Igor Dmochowski explained that LLMs are no longer limited just by model size or benchmark scores; the cost of longer reasoning means runtime, context length, memory use, and throughput matter just as much.</p>
<p>The post <a href="https://shiftmag.dev/llms-have-a-new-limit-it-costs-more-to-think-longer-12083/">LLMs Have a New Limit &#8211; It Costs More to Think Longer</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Modern agents use tools, handle text, images, and audio, remember context, use computers, and delegate tasks to subagents. But remember, <strong>each step adds more context, more tokens, and more compute</strong>. </p>



<p class="wp-block-paragraph">For developers, that means model quality still matters, but benchmark scores no longer tell the whole story. As Igor Dmochowski (Developer Relations Manager, Nvidia) said at <a href="https://shiftmag.dev/tag/infobip-shift-2026/" target="_blank" rel="noreferrer noopener">Infobip Shift 2026</a>:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">An agent is far more complex than a 2022-style chatbot: it can use tools, handle multimodal context, retain memory, and spawn subagents.</p>
</blockquote>



<h2 class="wp-block-heading"><span id="hybrid-architectures-preserve-useful-state-without-making-every-turn-more-expensive"><strong>Hybrid architectures preserve useful state without making every turn more expensive</strong></span></h2>



<p class="wp-block-paragraph">Longer context windows matter because agents need to remember more: past messages, documents, tool outputs, code, plans, and earlier steps. But as the context gets longer, self-attention becomes more expensive and harder to scale. </p>



<p class="wp-block-paragraph">Dmochowski’s point was that context length is a <strong>real systems constraint</strong>. A model may accept a huge prompt, but still be too slow or too memory-hungry for a real application. The question now is what they cost to process:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">The compute required to process your context was growing quadratically with regards to the context. This is not scalable.</p>
</blockquote>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="538" src="https://shiftmag.dev/wp-content/uploads/2026/09/igor_2-1-1024x538.png?x32039" alt="" class="wp-image-12136" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/igor_2-1-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/igor_2-1-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/igor_2-1-768x403.png 768w, https://shiftmag.dev/wp-content/uploads/2026/09/igor_2-1.png 1200w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Photo: Petar Krajičić Vilović</figcaption></figure>



<p class="wp-block-paragraph">The approach discussed on stage was a hybrid model that combines Transformer attention with Mamba-style state-space layers. Instead of using full attention for every token, the model uses <strong>more efficient layers for most of the work and keeps attention where it matters most</strong>. </p>



<p class="wp-block-paragraph">That matters for agents because long context often contains information with very different levels of importance. &#8220;A production system needs enough capacity to preserve useful state without making every turn proportionally more expensive. Hybrid architectures are one attempt to move that curve in a more manageable direction,&#8221; Dmochowski said.</p>



<h2 class="wp-block-heading"><span id="models-like-moe-active-compute-and-throughput-matter-more-than-total-parameter-count">Models like MoE, active compute and throughput matter more than total parameter count</span></h2>



<p class="wp-block-paragraph">The second shift is sparsity. </p>



<p class="wp-block-paragraph">In a conventional dense model, every token passes through the same large set of parameters. Igor explained that <a href="https://www.nvidia.com/en-us/glossary/mixture-of-experts/" target="_blank" rel="noreferrer noopener">Mixture of Experts</a>, or MoE, adds a routing mechanism <strong>that activates only a subset of specialized expert blocks for each token</strong>. The total model can contain far more parameters than are actually used during a single forward pass.</p>



<p class="wp-block-paragraph">According to Dmochowski, this distinction between total parameters and active parameters is becoming much more useful when comparing models. A large parameter count can suggest capacity, but it says much less about serving cost once sparse architectures enter the picture:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">We don’t need to put all of the tokens into all of the parameters of the model, we route them to the experts. For deployment, active compute, memory footprint, context length, KV cache behavior and tokens per second can be more informative than the headline parameter number alone.</p>
</blockquote>



<p class="wp-block-paragraph">Dmochowski connected that efficiency directly to capability. His point was that<strong> more throughput gives a system more room to reason within the same compute budget</strong>.</p>



<p class="wp-block-paragraph">Agents can do more intermediate work and finish long tasks with lower latency. When reasoning uses thousands of tokens, throughput becomes part of the capability envelope, which is why multi-token prediction matters.</p>



<p class="wp-block-paragraph">Lower-precision formats can reduce memory and compute requirements as well. Both techniques point toward the same goal: <strong>spend less hardware effort per useful token without destroying model quality</strong>:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">The faster model is going to be the smarter model. There is obviously a trade-off, but producing two or four times more tokens with the same compute gives you room to make that trade-off.</p>
</blockquote>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="538" src="https://shiftmag.dev/wp-content/uploads/2026/09/igor_3-1024x538.png?x32039" alt="" class="wp-image-12137" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/igor_3-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/igor_3-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/igor_3-768x403.png 768w, https://shiftmag.dev/wp-content/uploads/2026/09/igor_3.png 1200w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Photo: Petar Krajičić Vilović</figcaption></figure>



<h2 class="wp-block-heading"><span id="for-developers-efficiency-is-becoming-an-architectural-feature">For developers, efficiency is becoming an architectural feature</span></h2>



<p class="wp-block-paragraph">A model’s behavior <strong>depends on where you run it</strong> &#8211; the same model can act differently through a basic API than it does inside an agent harness with tools, memory, system instructions, and a loop that shows it the results of its actions. </p>



<p class="wp-block-paragraph">A leaderboard or prompt benchmark only tells part of the story. An agentic app is a full system, not just a model. &#8220;Tool schemas, retry logic, context construction, memory strategy and the execution loop all influence the final result. Evaluation should therefore resemble the environment in which the model will actually run,&#8221; Dmochowski said:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">The environment in which the models are set matters quite a lot. The behavior you get through an API can be completely different from the one you get within a harness.</p>
</blockquote>



<p class="wp-block-paragraph">The same idea applies during training and post-training: <strong>models should be trained in environments that simulate agent harnesses</strong>, and model choice and system design should be tested together.</p>



<h2 class="wp-block-heading"><span id="open-models-can-be-the-better-choice-when-private-data-has-to-stay-in-house">Open models can be the better choice when private data has to stay in-house</span></h2>



<p class="wp-block-paragraph">Dmochowski also made a case for open models in domains where the most valuable data cannot be handed to a frontier model provider. Internal code, proprietary documents, regulated records and specialized workflows may require organizations to adapt models within their own environment:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Open weights, training recipes and reproducible tooling can matter most in exactly those cases, because the competitive advantage is often the private data and workflow rather than the base model itself.</p>
</blockquote>



<p class="wp-block-paragraph">The useful takeaway from Dmochowski’s talk is that <strong>the bottlenecks are moving</strong>. The early LLM conversation focused on scale and benchmark scores, but agentic systems force developers to think about long context, memory, generation speed, tool use, and the cost of every extra reasoning step. That changes how models should be evaluated: a smaller or sparser model can be the better choice if it keeps quality while lowering latency, and a huge context window only matters if the system can afford to use it.</p>


<figure class="wp-block-post-featured-image"><img loading="lazy" decoding="async" width="1200" height="630" src="https://shiftmag.dev/wp-content/uploads/2026/09/igor_1.png?x32039" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="" style="object-fit:cover;" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/igor_1.png 1200w, https://shiftmag.dev/wp-content/uploads/2026/09/igor_1-768x403.png 768w, https://shiftmag.dev/wp-content/uploads/2026/09/igor_1-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/igor_1-1024x538.png 1024w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></figure><p>The post <a href="https://shiftmag.dev/llms-have-a-new-limit-it-costs-more-to-think-longer-12083/">LLMs Have a New Limit &#8211; It Costs More to Think Longer</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Only 7% of Organizations Using Kubernetes for AI Do So Daily</title>
		<link>https://shiftmag.dev/only-7-of-organizations-using-kubernetes-for-ai-do-so-daily-12086/</link>
		
		<dc:creator><![CDATA[Ivan Pelivanovic]]></dc:creator>
		<pubDate>Tue, 15 Sep 2026 11:35:26 +0000</pubDate>
				<category><![CDATA[DevOps]]></category>
		<category><![CDATA[Event]]></category>
		<category><![CDATA[Infobip Shift 2026]]></category>
		<category><![CDATA[Kubernetes]]></category>
		<guid isPermaLink="false">https://shiftmag.dev/?p=12086</guid>

					<description><![CDATA[<p>Two-thirds of organizations run AI workloads on Kubernetes, but only 7% do it daily. Katie Gamanji says the ecosystem still needs better operational maturity and standards.</p>
<p>The post <a href="https://shiftmag.dev/only-7-of-organizations-using-kubernetes-for-ai-do-so-daily-12086/">Only 7% of Organizations Using Kubernetes for AI Do So Daily</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></description>
										<content:encoded><![CDATA[<figure class="wp-block-post-featured-image"><img loading="lazy" decoding="async" width="1200" height="630" src="https://shiftmag.dev/wp-content/uploads/2026/09/katie_1.png?x32039" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="" style="object-fit:cover;" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/katie_1.png 1200w, https://shiftmag.dev/wp-content/uploads/2026/09/katie_1-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/katie_1-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/katie_1-768x403.png 768w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></figure>


<p class="wp-block-paragraph">Yesterday at <a href="https://shiftmag.dev/tag/infobip-shift-2026/" target="_blank" rel="noreferrer noopener">Infobip Shift 2026</a>, I listened to Apple Principal Engineer and CNCF Technical Oversight Committee member Katie Gamanji talk about the <strong>state of cloud native and its slow turn towards AI</strong>.</p>



<p class="wp-block-paragraph">She said cloud native spent years becoming stable and unexciting, and AI teams now need that stability more than new tools. The big question is whether the people around that infrastructure can move as fast as the platform itself.</p>



<p class="wp-block-paragraph">For teams already using Kubernetes in production, the next change is very concrete, Katie said:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Dynamic Resource Allocation, or DRA, reached general availability in Kubernetes 1.34, adding a native way to request GPU and TPU capacity for AI workloads.</p>
</blockquote>



<h2 class="wp-block-heading"><span id="kubernetes-is-now-officially-boring">Kubernetes is now officially boring</span></h2>



<p class="wp-block-paragraph">Katie started with a number that captures how far Kubernetes has come: &#8220;98% of organizations in the CNCF’s annual report said they had adopted cloud native technologies.&#8221; After a decade of development, Kubernetes has become mainstream:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Kubernetes is now described as being boring, and I think this is a wonderful achievement.</p>
</blockquote>



<p class="wp-block-paragraph">The problem appears when AI workloads enter that mature infrastructure. According to the figures Katie presented, <strong>66% of organizations are already deploying AI workloads on Kubernetes</strong>. But only 7% are doing it daily, while 47% deploy them occasionally. Of the organizations using Kubernetes for AI, 23% have fully adopted the Kubernetes stack and 43% have partially adopted it.</p>



<p class="wp-block-paragraph">That 67 to 7 gap points to one likeliest cause. The talk did not name the cause directly, but the pattern Katie described points to operational immaturity as the likeliest one:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Teams fine-tune an existing model rather than build one, deploy it once to prove it works, and lack the automated retraining pipelines that would treat the model as a dynamic component.</p>
</blockquote>



<p class="wp-block-paragraph">The 7% deploying daily, she noted, &#8220;most likely implemented an automated retraining pipeline that treats the model as a dynamic component rather than a static element.&#8221; <strong>Running an AI workload is not as simple as putting another application into a container</strong>. Teams need model artifacts, rollback strategies, and training pipelines. Kubernetes is being positioned as the shared platform, though whether teams adopt it that way is another question.</p>



<h2 class="wp-block-heading"><span id="dra-and-ai-conformance-are-making-kubernetes-the-ai-platform">DRA and AI Conformance are making Kubernetes the AI platform</span></h2>



<p class="wp-block-paragraph">Katie compared this to DevOps: Kubernetes once brought developers and operations together, and now it could do the same for infrastructure teams and data scientists. But the analogy has limits &#8211; DevOps took years to work, and it is still not fully settled everywhere.</p>



<p class="wp-block-paragraph"><strong>Data scientists</strong>, for their part, <strong>rarely want to operate Kubernetes</strong>; they want a model served. Still, a platform that can run both traditional workloads and training or inference jobs reduces the need for separate AI infrastructure. </p>



<p class="wp-block-paragraph">Kubernetes is also adapting for this: DRA in 1.34 gives teams a more reliable way to allocate accelerators, which matters most for teams deploying daily.</p>



<p class="wp-block-paragraph">The <a href="https://www.cncf.io/announcements/2025/11/11/cncf-launches-certified-kubernetes-ai-conformance-program-to-standardize-ai-workloads-on-kubernetes/" target="_blank" rel="noreferrer noopener">AI Conformance Working Group</a> defines what a platform must support to run AI workloads reliably, with portability as the goal as Katie said: </p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">If you have two conformant platforms, it&#8217;s going to be easy to lift and shift one product to the next platform.</p>
</blockquote>



<p class="wp-block-paragraph">Other initiatives target batch workloads, inference performance, and AI integration, and an Agent Sandbox effort is designing stateful, isolated runtimes for AI agents. The Serving Working Group completed its milestones and archived itself in February, continuing as a SIG on inference performance.</p>



<p class="wp-block-paragraph">The list is still changing quickly. Katie expects much of this landscape to look different within six months to a year.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="538" src="https://shiftmag.dev/wp-content/uploads/2026/09/katie_2-1024x538.png?x32039" alt="" class="wp-image-12131" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/katie_2-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/katie_2-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/katie_2-768x403.png 768w, https://shiftmag.dev/wp-content/uploads/2026/09/katie_2.png 1200w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption"> Photo: Petar Krajičić Vilović</figcaption></figure>



<h2 class="wp-block-heading"><span id="the-cloud-native-playbook-ai-can-borrow">The cloud native playbook AI can borrow</span></h2>



<p class="wp-block-paragraph">Katie divided the emerging open source AI ecosystem into three areas: <strong>training, inference, and agents</strong>. Training turns data into a model. Inference serves it. Agents connect it to the outside world.</p>



<p class="wp-block-paragraph">For platform leads, maturity matters. Training is led by the <a href="https://pytorch.org/foundation/" target="_blank" rel="noreferrer noopener">PyTorch Foundation</a>, inference has many options, and agents are still young under the <a href="https://aaif.io/" target="_blank" rel="noreferrer noopener">Agentic AI Foundation</a>.</p>



<p class="wp-block-paragraph"><strong>Cloud native experience can help AI tools mature faster</strong>. Security, observability, and identity were already solved in cloud native, and the CNCF project pipeline shows that kind of progress at scale.e.</p>



<p class="wp-block-paragraph">It also prunes: 28 archived projects. For anyone choosing a stack, the archive list is a practical filter, and Katie argued archival is a healthy sign, letting maintainers redirect energy toward projects that earn adoption.</p>



<h2 class="wp-block-heading"><span id="teams-that-engage-now-will-shape-the-patterns-everyone-else-follows">Teams that engage now will shape the patterns everyone else follows</span></h2>



<p class="wp-block-paragraph">For Katie, the next phase of cloud native is about whether the people building and maintaining the ecosystem can keep pace. She pointed out that contributing also means production feedback, feature requests, documentation, and white papers all shape projects.</p>



<p class="wp-block-paragraph">But in the end Katie&#8217;s talk left the hard questions unanswered, but the data shows why:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">98% of organizations trust Kubernetes with their infrastructure, two thirds have tried running AI on it, and almost nobody operates it continuously.</p>
</blockquote>



<p class="wp-block-paragraph">The specific bottlenecks, daily operational maturity, DRA adoption, and the conformance baseline, are still being defined, which means teams deploying AI on Kubernetes today are writing the patterns everyone else will copy:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">All of the working groups&#8217; meeting notes, invites, and repositories are public on GitHub. The teams that engage now, before the patterns harden, will not have to retrofit someone else&#8217;s choices in two years.</p>
</blockquote>



<p class="wp-block-paragraph"></p>
<p>The post <a href="https://shiftmag.dev/only-7-of-organizations-using-kubernetes-for-ai-do-so-daily-12086/">Only 7% of Organizations Using Kubernetes for AI Do So Daily</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>AI speeds up prototyping, but teams still need to distinguish decisions from assumptions</title>
		<link>https://shiftmag.dev/ai-speeds-up-prototyping-but-teams-still-need-to-distinguish-decisions-from-assumptions-12049/</link>
		
		<dc:creator><![CDATA[Anastasija Uspenski]]></dc:creator>
		<pubDate>Mon, 14 Sep 2026 14:15:01 +0000</pubDate>
				<category><![CDATA[Event]]></category>
		<category><![CDATA[Figma]]></category>
		<category><![CDATA[Infobip Shift 2026]]></category>
		<guid isPermaLink="false">https://shiftmag.dev/?p=12049</guid>

					<description><![CDATA[<p>Figma’s Designer Advocate argued that AI can make prototypes look finished too early, which can hide dangerous assumptions.</p>
<p>The post <a href="https://shiftmag.dev/ai-speeds-up-prototyping-but-teams-still-need-to-distinguish-decisions-from-assumptions-12049/">AI speeds up prototyping, but teams still need to distinguish decisions from assumptions</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></description>
										<content:encoded><![CDATA[<figure class="wp-block-post-featured-image"><img loading="lazy" decoding="async" width="1200" height="630" src="https://shiftmag.dev/wp-content/uploads/2026/09/hugo_3.png?x32039" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="" style="object-fit:cover;" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/hugo_3.png 1200w, https://shiftmag.dev/wp-content/uploads/2026/09/hugo_3-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/hugo_3-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/hugo_3-768x403.png 768w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></figure>


<p class="wp-block-paragraph">The easier it is to make something look finished, the more important it is to understand <strong>what the team actually decided</strong>.</p>



<p class="wp-block-paragraph">That was the starting point for Hugo Raymond, Designer Advocate at Figma, in his talk at <a href="https://shiftmag.dev/tag/infobip-shift-2026/" target="_blank" rel="noreferrer noopener">Infobip Shift 2026</a>.<br></p>



<p class="wp-block-paragraph">His main point was that <strong>AI becomes a problem when it fills in the gaps before decisions are made</strong>, because it can hide unresolved decisions.</p>



<h2 class="wp-block-heading"><span id="a-prototype-shouldn%e2%80%99t-try-to-look-like-a-finished-product"><strong>A prototype shouldn’t try to look like a finished product</strong></span></h2>



<p class="wp-block-paragraph">Hugo reminded us that the purpose of a prototype is to <strong>help the team test an idea</strong>. He compared it to a model airplane:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">A model doesn’t try to be a real airplane. We use it to study specific characteristics of the airplane.</p>
</blockquote>



<p class="wp-block-paragraph">The problem, he argued, is that AI can now create a very convincing prototype even when the original idea is still incomplete:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">In the past, gaps were easier to see. If a screen did not exist, a flow ended in a dead end, or the team had not defined a certain state, you could clearly see that nobody had made that decision yet. AI can now connect the flow, generate data, create loading and error states, and add behavior that nobody explicitly defined.</p>
</blockquote>



<p class="wp-block-paragraph">In the end, everything looks like one coherent product, even though the team designed some parts and the model simply filled in the rest.</p>



<h2 class="wp-block-heading"><span id="visual-polish-no-longer-proves-that-the-team-solved-the-problem">Visual polish no longer proves that the team solved the problem</span></h2>



<p class="wp-block-paragraph">According to Hugo, this creates an especially important problem for developers: <strong>the more finished a prototype looks, the more authority we give it</strong> and the more likely we are to treat it as a specification that developers simply need to implement.</p>



<p class="wp-block-paragraph">Hugo warned that a polished prototype no longer means that someone has thought through every edge case, business rule, or technical constraint.</p>



<p class="wp-block-paragraph">One of his key points was that <strong>polish used to be expensive</strong>. If something looked highly refined, you could at least assume that someone had spent time thinking through the problem. Now, the tables had turned:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">AI has broken that connection. The first draft can now look like the final iteration. Teams therefore need to make a much clearer distinction between what they have actually decided and what is still only an assumption.</p>
</blockquote>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="538" src="https://shiftmag.dev/wp-content/uploads/2026/09/hugo_1-1024x538.png?x32039" alt="" class="wp-image-12062" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/hugo_1-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/hugo_1-768x403.png 768w, https://shiftmag.dev/wp-content/uploads/2026/09/hugo_1-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/hugo_1.png 1200w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Photo: Petar Krajičić Vilović</figcaption></figure>



<h2 class="wp-block-heading"><span id="code-is-not-the-opposite-of-design">Code is not the opposite of design</span></h2>



<p class="wp-block-paragraph">Hugo doesn’t think the solution is to <strong>move design back into static mockups</strong>, quite the opposite: code is not the opposite of design but one of the materials we use to design:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">A code prototype can use real data, react to unexpected input, and show states that a designer did not draw in advance. AI has made this kind of experimentation much more accessible.</p>
</blockquote>



<p class="wp-block-paragraph">Still, <strong>code and a design canvas don&#8217;t show the same things</strong>.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Code is very good at showing what happens, while the canvas can do a better job of preserving why the team made a certain decision, which alternatives the team considered, and what the team still needs to solve.</p>
</blockquote>



<p class="wp-block-paragraph">Because of that, he sees a future where <strong>teams combine different materials depending on what they’re trying to learn.</strong></p>



<h2 class="wp-block-heading"><span id="ai-speeds-up-execution-but-not-decision-making">AI speeds up execution, but not decision-making</span></h2>



<p class="wp-block-paragraph">The final and most important point of his talk focused on the limits of automation.</p>



<p class="wp-block-paragraph">AI can take over a lot of repetitive work, speed up prototyping, and help teams test ideas more cheaply, but <strong>making something faster does not mean we understand what to make any better</strong>.</p>



<p class="wp-block-paragraph">Hugo therefore makes a distinction between friction worth removing and friction that serves a purpose:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">We can automate copying, waiting, and repetition. We should not automate comparing options, testing assumptions, discussing trade-offs, and questioning decisions.</p>
</blockquote>



<p class="wp-block-paragraph">As he pointed out, these moments help teams develop the judgment that AI cannot simply replace. “When almost every prototype can look finished, how convincing it looks is no longer the most important thing,” he said. What matters is whether the team can still clearly see what it has decided, what it has assumed, and what it still needs to solve.</p>
<p>The post <a href="https://shiftmag.dev/ai-speeds-up-prototyping-but-teams-still-need-to-distinguish-decisions-from-assumptions-12049/">AI speeds up prototyping, but teams still need to distinguish decisions from assumptions</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Google DeepMind: AI’s Next Step Is Learning From Its Own Experience</title>
		<link>https://shiftmag.dev/google-deepmind-ais-next-step-is-learning-from-its-own-experience-not-just-human-data-12023/</link>
		
		<dc:creator><![CDATA[Nikolina Oršulić]]></dc:creator>
		<pubDate>Mon, 14 Sep 2026 12:14:23 +0000</pubDate>
				<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Event]]></category>
		<category><![CDATA[Google]]></category>
		<category><![CDATA[Google DeepMind]]></category>
		<category><![CDATA[Infobip Shift 2026]]></category>
		<guid isPermaLink="false">https://shiftmag.dev/?p=12023</guid>

					<description><![CDATA[<p>We’re waiting for the next big AI leap - systems learning from their own experience rather than just human data - and at Infobip Shift, Google DeepMind’s Benoit Schillings said that’s exactly where the field is headed.</p>
<p>The post <a href="https://shiftmag.dev/google-deepmind-ais-next-step-is-learning-from-its-own-experience-not-just-human-data-12023/">Google DeepMind: AI’s Next Step Is Learning From Its Own Experience</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></description>
										<content:encoded><![CDATA[<figure class="wp-block-post-featured-image"><img loading="lazy" decoding="async" width="1200" height="630" src="https://shiftmag.dev/wp-content/uploads/2026/09/benoa_2_.png?x32039" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="" style="object-fit:cover;" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/benoa_2_.png 1200w, https://shiftmag.dev/wp-content/uploads/2026/09/benoa_2_-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/benoa_2_-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/benoa_2_-768x403.png 768w" sizes="auto, (max-width: 1200px) 100vw, 1200px" /></figure>


<p class="wp-block-paragraph">Google DeepMind’s VP of Technology argued that one of the most important developments to watch is the rise of systems that can <strong>learn beyond the limits of human-generated training data</strong>. </p>



<p class="wp-block-paragraph">He said AI has already gone through most publicly available human knowledge and is ready for the next step:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">What it really needs now is to start generating its own experience and its own data. In the next year or two, that’s exactly where AI is heading.<br></p>
</blockquote>



<p class="wp-block-paragraph">That’s where <strong>recursive self-improvement</strong> comes in: AI could start learning from its own experience by trying problems, checking its answers, finding mistakes, and testing different approaches instead of depending only on human-generated data. Humans do this naturally, but AI still struggles with it for now.</p>



<h2 class="wp-block-heading"><span id="ai-can-learn-from-experience-but-monitoring-has-to-improve-too"><strong>AI can learn from experience, but monitoring has to improve too</strong></span></h2>



<p class="wp-block-paragraph">Schillings pointed to <a href="https://deepmind.google/research/alphago/" target="_blank" rel="noreferrer noopener">AlphaGo</a> as an early example. </p>



<p class="wp-block-paragraph">AlphaGo is DeepMind’s system that beat one of the world’s best Go players, and Go is an ancient Chinese board game where the goal is to control territory on a grid. <strong>What made AlphaGo important was how it learned</strong>: by playing against itself and finding strategies that didn’t come from human games.</p>



<p class="wp-block-paragraph">Humans spent centuries developing Go, and now they study strategies discovered by AI. So: who is teaching whom? AI and humans learn from each other.</p>



<p class="wp-block-paragraph">There is, however, a major problem with the idea of AI learning from itself. What happens when the AI is wrong and who determines whether the data it produces is useful? This is particularly important because, as Schillings noted during his talk, models can sometimes <strong>find ways to appear successful without actually solving the problem</strong> they were given.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Cheating is a real problem. You ask a model to solve a problem, especially in code, and it will tell you, &#8220;I’ve done it, it’s there, it’s beautiful.&#8221; Then you check and find that it actually did not solve the problem. It may have stolen the result from somewhere else, or, even more amusingly, you ask it to write a tool for Linux and it says, &#8220;I wrote the tool,&#8221; when it’s really just invoking the existing tool and hiding the trace.</p>
</blockquote>



<p class="wp-block-paragraph"><strong>Sophisticated cheeting</strong> is becoming&nbsp;more and more&nbsp;an issue even for model to self-verify. So, we need to get the monitor to become more sophisticated also. That&#8217;s&nbsp;a part of&nbsp;the recursive&nbsp;self-improvement, he said.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="538" src="https://shiftmag.dev/wp-content/uploads/2026/09/benoa_3_-1024x538.png?x32039" alt="" class="wp-image-12033" srcset="https://shiftmag.dev/wp-content/uploads/2026/09/benoa_3_-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/09/benoa_3_-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/09/benoa_3_-768x403.png 768w, https://shiftmag.dev/wp-content/uploads/2026/09/benoa_3_.png 1200w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Photo: Petar Krajičić Vilović</figcaption></figure>



<h2 class="wp-block-heading"><span id="ai-needs-step-by-step-feedback">AI needs step-by-step feedback</span></h2>



<p class="wp-block-paragraph">Another part of Schillings’ point is <strong>how AI systems are trained</strong>. Some problems are easy to check: a proof works or it doesn’t, code passes tests or it doesn’t. But real-world problems are usually messier than that.</p>



<p class="wp-block-paragraph">Schillings compared it to teaching someone to climb a mountain with only one bit of feedback: you fell or you didn’t. That would be a terrible way to learn. A better teacher would give step-by-step feedback:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">This part was good. Try this section differently. You&#8217;re getting closer. Here&#8217;s where you can improve.</p>
</blockquote>



<p class="wp-block-paragraph"><strong>AI systems are being trained to give this kind of step-by-step reward</strong>. That matters more than it first seems: if models can learn not just from success and failure, but from how close they were to a better solution, they can handle much more complex problems.</p>



<h2 class="wp-block-heading"><span id="new-systems-may-reveal-things-we-haven%e2%80%99t-noticed-yet"><strong>New systems may reveal things we haven’t noticed yet</strong></span></h2>



<p class="wp-block-paragraph">What happens when AI starts exploring the unknown? This is where the implications become interesting. </p>



<p class="wp-block-paragraph">Code is a relatively convenient environment for AI because there is usually a way to test whether something works. Schillings expects the coming wave of AI development to produce an <strong>increasing number of breakthroughs</strong> in areas such as physical sciences, biology, mathematics and engineering. The reason is straightforward.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">If AI can generate hypotheses, test them, reject unsuccessful approaches and build on successful ones, it can start exploring the space of possibilities itself. That is a very different proposition from using AI as a better search engine. The ultimate value of these systems may be that they help us discover things nobody knows yet.</p>
</blockquote>



<p class="wp-block-paragraph">And that could turn out to be the most important AI story of the next few years.</p>
<p>The post <a href="https://shiftmag.dev/google-deepmind-ais-next-step-is-learning-from-its-own-experience-not-just-human-data-12023/">Google DeepMind: AI’s Next Step Is Learning From Its Own Experience</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>

<!--
Performance optimized by W3 Total Cache. Learn more: https://www.boldgrid.com/w3-total-cache/?utm_source=w3tc&utm_medium=footer_comment&utm_campaign=free_plugin

Page Caching using Disk: Enhanced 

Served from: shiftmag.dev @ 2026-10-06 22:33:11 by W3 Total Cache
-->