<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Adna Šestić, Author at ShiftMag</title>
	<atom:link href="https://shiftmag.dev/author/adna-sestic/feed/" rel="self" type="application/rss+xml" />
	<link>https://shiftmag.dev/author/adna-sestic/</link>
	<description>Insightful engineering content &#38; community</description>
	<lastBuildDate>Wed, 07 Oct 2026 13:21:46 +0000</lastBuildDate>
	<language>en-GB</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://shiftmag.dev/wp-content/uploads/2026/09/cropped-Shift-32x32.webp</url>
	<title>Adna Šestić, Author at ShiftMag</title>
	<link>https://shiftmag.dev/author/adna-sestic/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>How We Catch Messaging Campaign Drift Before It Spreads</title>
		<link>https://shiftmag.dev/how-we-catch-messaging-campaign-drift-before-it-spreads-11136/</link>
		
		<dc:creator><![CDATA[Adna Šestić]]></dc:creator>
		<pubDate>Wed, 07 Oct 2026 13:21:46 +0000</pubDate>
				<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Machine Learning]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[security]]></category>
		<guid isPermaLink="false">https://shiftmag.dev/?p=11136</guid>

					<description><![CDATA[<p>Over time, the content a messaging campaign sends can shift away from what was originally approved. We call this campaign drift, and in business messaging, it's a real and recurring problem.</p>
<p>The post <a href="https://shiftmag.dev/how-we-catch-messaging-campaign-drift-before-it-spreads-11136/">How We Catch Messaging Campaign Drift Before It Spreads</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></description>
										<content:encoded><![CDATA[<figure class="wp-block-post-featured-image"><img fetchpriority="high" decoding="async" width="1200" height="630" src="https://shiftmag.dev/wp-content/uploads/2026/08/ShiftMag-article-banner-1.png?x32039" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="" style="object-fit:cover;" srcset="https://shiftmag.dev/wp-content/uploads/2026/08/ShiftMag-article-banner-1.png 1200w, https://shiftmag.dev/wp-content/uploads/2026/08/ShiftMag-article-banner-1-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2026/08/ShiftMag-article-banner-1-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2026/08/ShiftMag-article-banner-1-768x403.png 768w" sizes="(max-width: 1200px) 100vw, 1200px" /></figure>


<p class="wp-block-paragraph">When businesses create messaging campaigns on Infobip, <strong>they must go through an approval process</strong>. They share brand details and campaign information, such as the purpose, target age group, and campaign type, which is checked for consistency and compliance with platform safety rules like WhatsApp or Viber terms of use.</p>



<p class="wp-block-paragraph">Once a campaign goes live, it’s our job as developers to ensure the messages being sent continue to reflect what was approved.</p>



<p class="wp-block-paragraph"><strong>There’s a whole spectrum of ways a campaign can drift</strong>, though. Some changes are obvious safety problems, like phishing or illegal content. Others are more subtle, like using the wrong brand voice, adding marketing language to transactional messages, or changing personalized content outside the approved template. The hardest cases are the ones that look compliant at first but still blur the line.</p>



<p class="wp-block-paragraph">&#8230;and that&#8217;s precisely what makes this a tough engineering problem.</p>



<h2 class="wp-block-heading"><span id="keyword-matching-isn%e2%80%99t-enough-to-catch-message-drift">Keyword matching isn’t enough to catch message drift</span></h2>



<p class="wp-block-paragraph"><strong>Legitimate traffic and drifting traffic can look very similar</strong>, even when they mean different things. Rule-based systems and keyword matching can’t always tell the difference. Detecting drift requires understanding meaning, context, and who is sending the message.</p>



<p class="wp-block-paragraph">If that&#8217;s a little hard to picture in the abstract, here&#8217;s an example. It&#8217;s rarely this overt in practice, but it makes the shape of the problem obvious.</p>



<p class="wp-block-paragraph">Imagine you&#8217;re a long-time client of an eye clinic. Every year, like clockwork, you receive the same familiar birthday message from the same phone number:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Hello Sam, this is Clermont Family Eyecare, we want to wish you a happy birthday! <img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f60e.png" alt="😎" class="wp-smiley" style="height: 1em; max-height: 1em;" /><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f389.png" alt="🎉" class="wp-smiley" style="height: 1em; max-height: 1em;" /></p>
</blockquote>



<p class="wp-block-paragraph">Generic, but consistent &#8211; a small customer-care touch from a business you trust.</p>



<p class="wp-block-paragraph">This year, the message arrives slightly altered, from that same number:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Hello Sammy, this is For Your Eyes Only Eye Care, we want to wish you a very happy birthday!</p>
</blockquote>



<p class="wp-block-paragraph">The tone is friendly and the structure looks the same. But the brand name has changed, and the link is new. You tap it, and suddenly you’re on a page with explicit content, not eyecare. <strong>In one click, trust is broken</strong> &#8211; even though the message came from the same number and passed every superficial check.</p>



<p class="wp-block-paragraph">On average, this system, when completed, is meant to process just short of 300 million SMS and RCS messages daily (crazy, right?). So, the engineering question becomes: how do you detect drift semantically, in production, at that scale, without reasoning about every message individually? Let me walk you through how we got to an answer.</p>



<h2 class="wp-block-heading"><span id="the-system-needs-to-be-designed-from-first-principles">The system needs to be designed from first principles</span></h2>



<p class="wp-block-paragraph">A few observations about the traffic led to the architecture.</p>



<p class="wp-block-paragraph">The first is the most obvious: <strong>messaging campaigns are repetitive by nature</strong>. Messages are templated, sometimes identical within a campaign. At our scale, this immediately points toward grouping messages and reasoning about drift at the group level rather than per message.</p>



<p class="wp-block-paragraph">There’s an important operational detail here you should know: when businesses create messaging campaigns on our platform, they go through an approval process. They provide brand information and campaign metadata like intent, recipient age or campaign type categorization, which is reviewed for internal consistency and against platform safety guidelines (think WhatsApp or Viber Terms of Use). Once a campaign goes live, it’s our job to ensure the messages being sent continue to reflect what was approved.</p>



<p class="wp-block-paragraph">You can see that the problem isn’t really about tracking a semantic distribution of campaign messages and watching it shift over time, even though the word drift might suggest that. Historical traffic isn’t what defines drift here. The problem is closer to <strong>continuous content validation</strong>: preserving the intent stated in the original campaign metadata and looking for new patterns that warrant separate interpretation and potential action.</p>



<p class="wp-block-paragraph">Even though this framing focuses on novel data,<strong> historical traffic still matters</strong>. If incoming messages reliably match traffic that has already been evaluated, we can decide in real time whether they’re legitimate or drifting. Messages that match approved intent should be handled immediately, and only unexplained content needs to be isolated for further analysis.</p>



<p class="wp-block-paragraph">A message entering our system takes one of two paths:</p>



<ul class="wp-block-list">
<li>A real-time <strong>classification path</strong>, which assigns it to a previously observed group with a known verdict, or</li>



<li>A deferred <strong>discovery path</strong>, where it’s held aside to help form a new group.</li>
</ul>



<p class="wp-block-paragraph">And there you have it: the gist of the design. The rest of the post is about how we turn that into a production system.</p>



<h2 class="wp-block-heading"><span id="the-architecture">The architecture</span></h2>



<p class="wp-block-paragraph">In practice, this is a multi-stage microservice pipeline. A message flows through it as follows:</p>



<ol class="wp-block-list">
<li><strong>Ingestion and reassembly</strong>.<strong> </strong>The orchestrator service ingests incoming traffic and reassembles multi-part messages into their original textual form.</li>



<li><strong>Vectorization</strong>. To compare messages by intent rather than wording, we map them into a shared semantic space. A dedicated vectorization service takes a batch of messages and returns their embedding vectors. Messages expressing similar intent end up close together in this space, even when the wording differs. That way, grouping by proximity in the embedding space gives us grouping by semantics.</li>



<li><strong>Search</strong>.<strong> </strong>The orchestrator queries an Elasticsearch index of known cluster representatives using the message’s embedding. If a match is found, the message is assigned to that cluster, and the cluster signature is promoted to a local cache, so subsequent traffic from the same sender resolves faster. This is the classification path in action.</li>



<li><strong>Noise accumulation</strong>.<strong> </strong>If no match is found, the message takes the discovery path: it’s written to a <em>noise</em> index. The name comes from clustering terminology: technically, it’s an “unclassified” index, holding candidate data for future clusters.</li>



<li><strong>Clustering</strong>.<strong> </strong>A scheduled job continuously pulls accumulated noise and runs clustering to discover new semantic groups. Clustering is performed on sender level. In practice, we’ve observed that for a sender, message volume is a great primary signal for parameter configuration, and after an obscene amount of per-sender experiments, we did manage to settle on a base of volume-keyed parameter sets. Although not as ubiquitously successful as maybe a best-of-n configs approach would have been, at our scale the compute savings make it the right call.</li>



<li><strong>LLM-as-a-judge</strong>.<strong> </strong>Finally, cluster representatives are relayed to an LLM, which produces a verdict for the <em>cluster</em>, not for individual messages. The model also outputs per-message reasoning, which is what makes the verdict actionable downstream.</li>
</ol>



<figure class="wp-block-image size-large"><img decoding="async" width="1024" height="614" src="https://shiftmag.dev/wp-content/uploads/2026/10/clustering_3200x1920-1024x614-1.jpg?x32039" alt="" class="wp-image-12434" srcset="https://shiftmag.dev/wp-content/uploads/2026/10/clustering_3200x1920-1024x614-1.jpg 1024w, https://shiftmag.dev/wp-content/uploads/2026/10/clustering_3200x1920-767x460.jpg 767w, https://shiftmag.dev/wp-content/uploads/2026/10/clustering_3200x1920-300x180.jpg 300w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<h2 class="wp-block-heading"><span id="walking-the-eyecare-message-through-the-pipeline">Walking the eyecare message through the pipeline</span></h2>



<p class="wp-block-paragraph">Let’s replay the message path from the start with this architecture in mind.</p>



<p class="wp-block-paragraph"><strong>The message arrives at ingestion</strong> alongside other traffic from the same eyecare clinic sender. The orchestrator reassembles it and the vectorization service returns its embedding. The orchestrator queries Elasticsearch for known clusters associated with this sender, and let’s imagine that there are several: the standard birthday message, an annual check-up reminder, a glasses-ready-for-pickup notification.</p>



<p class="wp-block-paragraph">The incoming embedding sits near the birthday cluster, but not near enough to match. The brand name swap and the shift from “Sam” to “Sammy” are subtle on paper, but they move the vector just far enough away.</p>



<p class="wp-block-paragraph">No match. The message goes to the noise index.</p>



<p class="wp-block-paragraph">Over the next clustering interval, <strong>more messages like it accumulate</strong> (you can be pretty sure they would, this reads like a classic sender takeover). The clustering job runs, a new cluster forms around them, and a handful of representatives are stored and passed to the LLM judge.</p>



<p class="wp-block-paragraph">The judge compares the cluster against the sender’s approved campaign metadata and <strong>flags it for brand impersonation</strong>. (The changed recipient personalization remains unflagged: on its own, it&#8217;s a minor surface variation that doesn&#8217;t shift the underlying intent.) Reasoning is attached. A human reviewer picks it up, and the sender is paused before the next batch of spam goes out.</p>



<p class="wp-block-paragraph">The message arrived through the same number, used the same structure, and would have passed any rule-based check. The only reason the system caught it was because the meaning had shifted.</p>



<figure class="wp-block-image size-large"><img decoding="async" width="1024" height="614" src="https://shiftmag.dev/wp-content/uploads/2026/10/message_path_3200x1920-1024x614.jpg?x32039" alt="" class="wp-image-12435" srcset="https://shiftmag.dev/wp-content/uploads/2026/10/message_path_3200x1920-1024x614.jpg 1024w, https://shiftmag.dev/wp-content/uploads/2026/10/message_path_3200x1920-300x180.jpg 300w, https://shiftmag.dev/wp-content/uploads/2026/10/message_path_3200x1920-767x460.jpg 767w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<h2 class="wp-block-heading"><span id="clustering-is-the-leverage-point">Clustering is the leverage point</span></h2>



<p class="wp-block-paragraph">At 300 million messages a day, per-message <strong>LLM inference is a non-starter on cost and latency alone</strong>. It would also be wasteful, because most incoming traffic is heavily templated or duplicated. Grouping by intent first means we make one judgment instead of a million.</p>



<p class="wp-block-paragraph">Embeddings group messages together, the cluster amortizes the cost of judging them, and the LLM is reserved for the moment judgment is actually necessary: when something new appears.</p>



<p class="wp-block-paragraph">Two (in my opinion) interesting<strong> design decisions</strong> follow from this:</p>



<ul class="wp-block-list">
<li>First, <strong>we deliberately don&#8217;t let clusters evolve</strong>. A cluster represents a stable behavior snapshot, a fixed meaning at a point in time. Letting clusters expand would let them quietly absorb new variations in intent and language, which is exactly what we&#8217;re trying to detect. Keeping them fixed forces any notable change in semantics to surface as a new cluster, and makes downstream judgment more coherent because the examples within a cluster genuinely belong together.</li>



<li>Second, <strong>this tightening of group intent allows us to represent each cluster by a few select examples rather than its full membership</strong>. This keeps the classification index lean for fast matching, and it gives the next stage, cluster judgment, a focused, reviewable input.</li>
</ul>



<h2 class="wp-block-heading"><span id="tradeoffs-we-accepted">Tradeoffs we accepted</span></h2>



<p class="wp-block-paragraph">As no architecture is without its flaws, the most interesting one is <strong>overclustering</strong>. </p>



<p class="wp-block-paragraph">The system inherits the limitations of its embedding model: embeddings capture intent well, but not every kind of shift. We expected greedy cluster expansion, but in practice the embedding space often treats benign surface variation (like a swapped greeting or different recipient name) as meaningfully different. Legitimate traffic from one campaign gets fragmented across many clusters, which inflates downstream cost and dilutes the signal we care about. </p>



<p class="wp-block-paragraph">The lesson is that embeddings alone won’t always get us to clean intent groups.</p>



<p class="wp-block-paragraph">The other two costs are more predictable:</p>



<ul class="wp-block-list">
<li><strong>Discovery latency. </strong>The classification path is real-time, but the discovery path isn&#8217;t: messages that don&#8217;t match an existing cluster sit in the noise index until the next clustering job runs. For genuinely novel drift, this means a delay between first appearance and first judgment. We tune the noise accumulation period to keep this window short for senders above a certain message volume, assuming that low-volume senders will continue to grow with time and graduate to a size above the threshold.</li>



<li><strong>Cluster proliferation from non-evolving clusters.</strong> Forcing semantic shifts to surface as new clusters means a sender with legitimate variation across message templates ends up represented by many clusters. We manage this with cluster aging policies explicitly, but the lean cluster representations also help in that regard, so the memory cost remains bounded.</li>
</ul>



<h2 class="wp-block-heading"><span id="model-provides-a-recommendation-not-a-decision">Model provides a recommendation, not a decision</span></h2>



<p class="wp-block-paragraph">A verdict on a cluster may end the pipeline, but <strong>the case is only closed when a human says it is</strong>.</p>



<p class="wp-block-paragraph">When the LLM judge flags a cluster, the verdict and reasoning go to a human review queue. Reviewers see the cluster representatives, the approved campaign metadata, and the model’s reasoning side by side. They can confirm the verdict, override it, or adjust its severity, and their decision propagates downstream.</p>



<p class="wp-block-paragraph">From there, <strong>action depends on severity</strong>: mild drift leads to client notification and guidance, more serious cases can block specific message types and trigger a root cause analysis, and the most severe, spam-grade drift can block senders entirely.</p>



<p class="wp-block-paragraph">The important thing to note is that the model provides a recommendation, not a decision.</p>



<h2 class="wp-block-heading"><span id="the-scariest-drift-is-the-one-that-still-looks-trusted">The scariest drift is the one that still looks trusted</span></h2>



<p class="wp-block-paragraph">A few things became clearer as we built this out than they were going in:</p>



<ul class="wp-block-list">
<li><strong>The most interesting drift is quiet</strong>: a message that looks familiar, comes from the same number, but means something slightly different and slowly erodes trust before anyone notices.</li>



<li><strong>Judging clusters instead of individual messages changes the question</strong>. <em>Is this message okay? </em>is narrow and easy to answer wrong. <em>What is this sender doing now?</em> is the question that matters, and a group of similar intent is the right unit to ask it about.</li>
</ul>



<p class="wp-block-paragraph">What’s immediately next is concrete, on two fronts.</p>



<ol class="wp-block-list">
<li>The first is <strong>resolving the overclustering issue</strong> described earlier: collapsing benign surface variation back into single clusters so the LLM judge isn&#8217;t doing redundant work.&nbsp; Lightweight named entity recognition and a smidge of the old regex to capture emails or phone numbers should do the trick, we hope.</li>



<li>The second is <strong>scaling</strong>: the system currently runs in a subset of regions and handles on the order of tens of millions of messages a week, and the immediate work is extending that toward the 300-million-a-day ceiling it&#8217;s designed for.</li>
</ol>



<p class="wp-block-paragraph">The eyecare message at the start of this post is just a small example of a bigger problem. What we’re really trying to defend is the meaning of a trusted channel.</p>



<p class="wp-block-paragraph">And for now, <strong>please don’t click weird links in your text</strong>s. Sam learned the hard way, and we’re not quite there yet on link-checking.</p>
<p>The post <a href="https://shiftmag.dev/how-we-catch-messaging-campaign-drift-before-it-spreads-11136/">How We Catch Messaging Campaign Drift Before It Spreads</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>

<!--
Performance optimized by W3 Total Cache. Learn more: https://www.boldgrid.com/w3-total-cache/?utm_source=w3tc&utm_medium=footer_comment&utm_campaign=free_plugin

Page Caching using Disk: Enhanced 

Served from: shiftmag.dev @ 2026-10-07 19:45:27 by W3 Total Cache
-->