<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>Yohangel Ramos’s blog</title>
    <link>https://yohangel.com/en/blog/</link>
    <atom:link href="https://yohangel.com/en/rss.xml" rel="self" type="application/rss+xml"/>
    <description>A builder’s notes: AI-assisted development, serverless architecture on AWS and modern frontend.</description>
    <language>en</language>
    <copyright>© 2026 Yohangel Ramos</copyright>
    <managingEditor>yohangelr@gmail.com (Yohangel Ramos)</managingEditor>
    <lastBuildDate>Tue, 18 Aug 2026 00:00:00 GMT</lastBuildDate>
    <image>
      <url>https://yohangel.com/icon-512.png</url>
      <title>Yohangel Ramos’s blog</title>
      <link>https://yohangel.com/en/blog/</link>
    </image>
    <item>
      <title>Lambda vs PostgreSQL: the day you run out of connections</title>
      <link>https://yohangel.com/en/blog/lambda-postgres-conexiones-rds-proxy/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/lambda-postgres-conexiones-rds-proxy/</guid>
      <description>A connection pool and an ephemeral function are two ideas that contradict each other. Why connections run out, what RDS Proxy actually fixes, and what you can fix for free before paying for it.</description>
      <pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>AWS</category>
      <content:encoded><![CDATA[<p><strong>The error always shows up at the worst possible time: <code>remaining connection slots are reserved for non-replication superuser connections</code>. The app had been running fine for weeks, nobody touched the database code, and then one traffic spike leaves half your API returning 500s. Everyone reaches for the same fix first: bump <code>max_connections</code> on the instance. It works for a while, and then it happens again. The number was never the problem. The problem is that a connection pool and an ephemeral function are two ideas that contradict each other.</strong></p>
<h2>One pool per environment, not one pool per application</h2>
<p>A pool exists to amortize something expensive. Opening a PostgreSQL connection costs real work: TCP handshake, TLS, authentication, and the server spins up a dedicated process per connection. In a long-lived backend — a Nest service on Fargate, say — you open ten connections at boot and reuse them for days. You pay that cost once.</p>
<p>Lambda has no &quot;at boot&quot; in the sense you&#39;re imagining. It has one boot per <strong>execution environment</strong>, and AWS creates as many environments as you have concurrent invocations. If your peak is 200 concurrent invocations, you have 200 independent Node processes, each with its own pool. A pool of 10 per environment means 2,000 connections knocking on a database that can handle 100.</p>
<p>The math is brutally simple:</p>
<pre><code>connections = Lambda concurrency × pool size per environment
</code></pre>
<p>A <code>db.t4g.medium</code> sits around 220 <code>max_connections</code>. At 50 concurrent invocations with a pool of 5, you&#39;re already at the ceiling. And unlike a server, you don&#39;t get to choose the concurrency here — traffic chooses it for you.</p>
<h2>The mistake that multiplies the damage: creating the pool inside the handler</h2>
<p>Before touching any infrastructure, look at where the client gets created. This is the pattern I run into most often, and it turns a scaling problem into an immediate one:</p>
<pre><code class="language-typescript">// BAD: a fresh pool on every invocation, and connections left dangling
export const handler = async (event) =&gt; {
  const pool = new Pool({ connectionString: process.env.DATABASE_URL });
  const { rows } = await pool.query(&#39;SELECT id FROM users WHERE email = $1&#39;, [event.email]);
  return rows[0];
};
</code></pre>
<p>Every invocation opens a brand new pool. And if the <code>end()</code> is missing — it almost always is — the connection stays occupied until PostgreSQL reaps it on timeout. You&#39;ve just turned every request into a leak.</p>
<p>The fix is to hoist it out of the handler, into module scope. That code runs once per execution environment, and the environment gets reused across invocations for as long as it stays warm:</p>
<pre><code class="language-typescript">// The pool lives in the execution environment and survives across invocations
const pool = new Pool({
  connectionString: process.env.DATABASE_URL,
  max: 1,                       // one environment serves one invocation at a time
  idleTimeoutMillis: 30000,     // release the connection when the environment goes idle
  connectionTimeoutMillis: 5000,
});

export const handler = async (event) =&gt; {
  const { rows } = await pool.query(&#39;SELECT id FROM users WHERE email = $1&#39;, [event.email]);
  return rows[0];
};
</code></pre>
<p><code>max: 1</code> looks wrong if you come from servers, but it&#39;s exactly right here: a Lambda environment handles <strong>one</strong> invocation at a time, so a pool of 10 is nine connections that never do any parallel work and absolutely do count against your limit.</p>
<h2>Reserved concurrency: the ceiling you actually control</h2>
<p>With <code>max: 1</code>, your connection count becomes precisely your concurrency. And concurrency you can cap, function by function:</p>
<pre><code class="language-hcl">resource &quot;aws_lambda_function&quot; &quot;api&quot; {
  function_name                  = &quot;api-handler&quot;
  reserved_concurrent_executions = 40   # hard ceiling of 40 connections
  # ...
}
</code></pre>
<p>It&#39;s an uncomfortable call, because it means accepting throttling: past 40 simultaneous invocations, Lambda starts rejecting. But I&#39;ll take a bounded failure in one function over a database that stops accepting connections and takes down the workers, the migrations and the admin panel along with it — none of which had anything to do with the spike. A limit isn&#39;t a restriction. It&#39;s deciding in advance who suffers first when something breaks.</p>
<h2>RDS Proxy: what it solves and what it doesn&#39;t</h2>
<p>RDS Proxy sits between Lambda and the database and maintains its own set of already-open connections. Your functions talk to the proxy, and the proxy multiplexes: many ephemeral clients over a handful of real, reused connections.</p>
<p>What you genuinely gain:</p>
<ul>
<li><strong>It absorbs spikes.</strong> When the database is at its ceiling, the proxy queues instead of rejecting.</li>
<li><strong>It removes connection setup cost</strong> from every cold start. That shows up in p99 latency more than you&#39;d expect.</li>
<li><strong>Cleaner failover.</strong> It holds the client connection open while the engine switches over.</li>
</ul>
<p>What&#39;s worth knowing before you adopt it:</p>
<ul>
<li><strong>It isn&#39;t free.</strong> It bills per vCPU of the database instance, and on small workloads it can cost more than simply sizing the database up.</li>
<li><strong>Pinning will ruin the party.</strong> Long transactions, session-level prepared statements, <code>SET</code> on session variables or temp tables all make the proxy pin a connection to your session and stop multiplexing. You end up paying for a proxy that reuses nothing. It&#39;s visible in CloudWatch as <code>DatabaseConnectionsCurrentlySessionPinned</code>, and it&#39;s the first thing I check when someone tells me the proxy &quot;did nothing&quot;.</li>
<li><strong>It lives inside the VPC.</strong> Which means putting your Lambda in private subnets, with everything that drags along in cold starts and networking bills.</li>
</ul>
<h2>What I&#39;d do today</h2>
<p>The order matters, and almost nobody follows it:</p>
<ol>
<li><strong>Hoist the pool out of the handler and set <code>max: 1</code>.</strong> Five minutes, zero cost, and it solves most cases outright.</li>
<li><strong>Set <code>reserved_concurrent_executions</code></strong> on the functions that touch the database. Also free, and it turns a total outage into a bounded degradation.</li>
<li><strong>Measure before you buy.</strong> Watch <code>DatabaseConnections</code> in CloudWatch during a real spike. If you&#39;re nowhere near the ceiling, you don&#39;t need a proxy.</li>
<li><strong>Add RDS Proxy</strong> when traffic is genuinely spiky and you&#39;ve confirmed your queries don&#39;t trigger pinning.</li>
</ol>
<p>And one non-technical note: if 90% of your application is CRUD against PostgreSQL with reasonably steady traffic, the right answer probably isn&#39;t any of those four. It&#39;s a long-lived container with a normal, boring pool. Lambda is superb for bursty work. Against a relational database under sustained load, you&#39;re often paying in complexity to solve a problem you never had.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Loading states: that centered spinner is a product decision</title>
      <link>https://yohangel.com/en/blog/estados-de-carga-skeletons-streaming/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/estados-de-carga-skeletons-streaming/</guid>
      <description>While the data loads there is a screen almost nobody designs. Skeletons that do not shift, streaming with Suspense, and the delay trick that makes an app feel fast.</description>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>Frontend</category>
      <content:encoded><![CDATA[<p><strong>There is one screen almost nobody designs and every user sees: the one that shows up while the data loads. The Figma mockup arrives with the data already in place, pretty and complete, and the loading state gets resolved in the last commit with an <code>if (loading) return &lt;Spinner /&gt;</code>. That if is a product decision made at eleven at night by whoever was in a hurry. And it shows: the app feels slow, the page jumps when the data lands, and the user cannot tell whether to wait or reload. I have spent years fixing this in dashboards, in a meal planner and in internal admin panels, and my conclusion is that loading states account for most of the speed your users think your product has.</strong></p>
<h2>The centered spinner is the worst possible default</h2>
<p>A spinner in the middle of the screen says exactly one thing: &quot;something is happening, I do not know what, I do not know how long&quot;. It wipes out the structure you already knew and replaces it with a wheel. When the data arrives, the interface appears all at once in a different position and your eye has to re-orient itself.</p>
<p>The underlying problem is that it treats the page as an atomic unit: either everything is there or nothing is. That is almost never true. In a typical dashboard, the title, the navigation, the filters and the shape of the cards are all known before you request a single byte from the server. The only missing part is the numbers. Blocking the 90% of the interface you could already paint in order to wait for the 10% you cannot is throwing away perceived performance.</p>
<p>The rule I follow: <strong>paint everything you already know, and mark as pending only what genuinely depends on the server.</strong></p>
<h2>Skeletons that do not shift: reserve the real space</h2>
<p>A skeleton is a placeholder shaped like the final content. It works for two reasons: it communicates what is about to appear and, more importantly, it reserves the space so nothing jumps when it does. If your skeleton is 40px tall and the real row is 72px, you have traded a spinner for a layout shift, which is worse.</p>
<p>That is why I build them from the same component instead of a parallel mockup that drifts out of sync on day one:</p>
<pre><code class="language-jsx">function CandidateRow({ data }) {
  return (
    &lt;li className=&quot;row&quot;&gt;
      &lt;span className=&quot;avatar&quot;&gt;{data ? &lt;img src={data.avatar} alt=&quot;&quot; /&gt; : null}&lt;/span&gt;
      &lt;span className=&quot;name&quot;&gt;{data ? data.name : &lt;Block w=&quot;60%&quot; /&gt;}&lt;/span&gt;
      &lt;span className=&quot;title&quot;&gt;{data ? data.title : &lt;Block w=&quot;35%&quot; /&gt;}&lt;/span&gt;
    &lt;/li&gt;
  );
}
</code></pre>
<p>The row keeps the same height and the same grid in both states, so going from skeleton to data is a content change, not a layout change. CSS does the heavy lifting:</p>
<pre><code class="language-css">.row { display: grid; grid-template-columns: 40px 1fr 1fr; min-height: 72px; }
.avatar { aspect-ratio: 1; border-radius: 50%; background: var(--gray-200); }

@media (prefers-reduced-motion: no-preference) {
  .block { animation: pulse 1.4s ease-in-out infinite; }
}
</code></pre>
<p>Two details that always get forgotten: <code>min-height</code> on the row (it prevents CLS when the real content is taller) and honoring <code>prefers-reduced-motion</code>, because a screen full of pulsing blocks is exactly the kind of animation that makes some people queasy.</p>
<h2>Streaming with Suspense: do not wait for the slowest call</h2>
<p>When a page requests four things, the usual shape is 80ms, 120ms, 150ms… and 900ms. If you wait for all of them before painting, your page takes 900ms. With streaming, the server sends the HTML for whatever is ready and the rest arrives later over the same connection.</p>
<p>In React you express this with <code>Suspense</code> boundaries. Each boundary is a promise to the user: &quot;this part is coming later, the rest is already here&quot;.</p>
<pre><code class="language-jsx">export default function Dashboard() {
  return (
    &lt;Layout&gt;
      &lt;Header /&gt;                          {/* instant, no data */}
      &lt;Suspense fallback={&lt;KpisSkeleton /&gt;}&gt;
        &lt;Kpis /&gt;                          {/* fast */}
      &lt;/Suspense&gt;
      &lt;Suspense fallback={&lt;TableSkeleton rows={8} /&gt;}&gt;
        &lt;ReportsTable /&gt;                  {/* the slow one: 900ms */}
      &lt;/Suspense&gt;
    &lt;/Layout&gt;
  );
}
</code></pre>
<p>The syntax is not the interesting part — where you put the boundaries is. One boundary per page buys you nothing: you are back to the global spinner with extra steps. One boundary per tiny component is not better either: you end up with a screen that flickers in pieces for two seconds and feels broken. My rule is to wrap <strong>blocks the user perceives as a single unit</strong> — a card, a table, a sidebar — and above all to isolate the slow thing so it does not hold the fast things hostage.</p>
<h2>The detail that killed the most complaints: delay and hold</h2>
<p>If a request takes 90ms and you show a skeleton, the user sees a flicker: it appears and disappears before the eye can process it. That flash reads as a glitch, not as speed. A skeleton shown for 40ms is more annoying than no skeleton at all.</p>
<p>The fix is not technical, it is about timing: <strong>do not show the loading state before ~200ms, and once you show it, hold it for at least ~400ms.</strong></p>
<pre><code class="language-ts">export function useVisibleLoading(loading: boolean) {
  const [visible, setVisible] = useState(false);

  useEffect(() =&gt; {
    if (!loading) return;
    const delay = setTimeout(() =&gt; setVisible(true), 200);
    return () =&gt; clearTimeout(delay);       // finished early: never seen
  }, [loading]);

  useEffect(() =&gt; {
    if (loading || !visible) return;
    const floor = setTimeout(() =&gt; setVisible(false), 400);
    return () =&gt; clearTimeout(floor);       // already visible: hold 400ms
  }, [loading, visible]);

  return visible;
}
</code></pre>
<p>With this, fast responses show nothing at all — which is what you want, since they are already fast — and slow ones show a stable state instead of a flicker. It is the smallest change I have shipped with the most direct impact on &quot;hey, the app feels way better now&quot;.</p>
<h2>What I do today on every new screen</h2>
<p>Before writing the first <code>fetch</code>, I ask three questions: what can I paint with no data (that goes outside every boundary), which blocks does the user read as units (that is where the Suspense boundaries and their skeletons go), and which call is the slow one (that one always gets isolated, even if it is the most important thing on the screen).</p>
<p>The mental mistake that took me years to unlearn was thinking of loading as a binary moment — loading or loaded — when it is really a sequence you get to choreograph. The same 900ms request can feel like a broken app or a living one, and the difference is not in the backend: it is in what you choose to show during those 900ms. Getting that call from 900ms to 600ms costs an afternoon and almost nobody notices; designing what people look at in the meantime costs an hour and everybody does.</p>
]]></content:encoded>
    </item>
    <item>
      <title>CloudFront in front of your app: why your hit rate is 30%</title>
      <link>https://yohangel.com/en/blog/cloudfront-cache-key-invalidaciones/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/cloudfront-cache-key-invalidaciones/</guid>
      <description>Adding a CDN is not caching. The cache key, your origin's Cache-Control and how you invalidate decide whether CloudFront saves you money or just adds one more hop.</description>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>AWS</category>
      <content:encoded><![CDATA[<p><strong>Putting CloudFront in front of an app is one of those decisions everyone applauds in the meeting and almost nobody verifies afterward. The distribution gets created, the domain points at it, the first test shows better latency, and everyone assumes it&#39;s &quot;cached now&quot;. Months later you open the metrics and the hit rate is 30%: seven out of ten requests still reach your origin, you&#39;re paying transfer twice, and you now have an extra layer to debug through when something looks odd. It&#39;s not that the CDN doesn&#39;t work. It&#39;s that caching isn&#39;t a checkbox: it&#39;s three decisions —what identifies a response, how long it lives, and how it gets replaced— and those three are yours, not AWS&#39;s.</strong></p>
<h2>The cache key is 90% of the problem</h2>
<p>A CDN stores responses indexed by a key. If two requests produce the same key, the second one is a hit. Everything you put into that key multiplies the number of possible entries and divides your hit rate.</p>
<p>The case I&#39;ve run into most: someone forwards the session cookie to the origin &quot;because the app needs it&quot; and accidentally puts it in the cache key too. From that moment on, every user gets their own copy of every resource. Technically there&#39;s a cache; in practice it caches nothing. Same story with forwarding all query strings: a campaign&#39;s <code>utm_source</code> values turn one URL into fifty.</p>
<p>The piece that fixes this is realizing CloudFront has two separate policies: the <strong>cache policy</strong> defines what goes into the key, and the <strong>origin request policy</strong> defines what gets sent to the origin. They are not the same thing. The cookie can travel to your origin without fragmenting the cache.</p>
<pre><code class="language-hcl">resource &quot;aws_cloudfront_cache_policy&quot; &quot;static&quot; {
  name        = &quot;static-assets&quot;
  min_ttl     = 0
  default_ttl = 86400
  max_ttl     = 31536000

  parameters_in_cache_key_and_forwarded_to_origin {
    enable_accept_encoding_brotli = true
    enable_accept_encoding_gzip   = true

    cookies_config { cookie_behavior = &quot;none&quot; }
    headers_config { header_behavior = &quot;none&quot; }

    query_strings_config {
      query_string_behavior = &quot;whitelist&quot;
      query_strings { items = [&quot;v&quot;] }
    }
  }
}
</code></pre>
<p>The rule I apply without exceptions: the cache key starts empty and you add whatever you can justify. Never the other way around.</p>
<h2><code>Cache-Control</code> decides; the distribution only sets limits</h2>
<p>The distribution&#39;s TTLs confuse a lot of people because they look like the main configuration, and they aren&#39;t. <code>default_ttl</code> only applies when the origin sends <strong>no</strong> cache headers. <code>max_ttl</code> is a ceiling. <code>min_ttl</code> is a floor. The real decision-maker is your origin, route by route.</p>
<p>And that&#39;s where the header with the best performance-per-character I know lives: <code>s-maxage</code>.</p>
<pre><code class="language-http"># Bundle with a content hash in its name: genuinely immutable
Cache-Control: public, max-age=31536000, immutable

# HTML for a page that changes: the browser doesn&#39;t keep it, the CDN does
Cache-Control: public, max-age=0, s-maxage=300, stale-while-revalidate=86400
</code></pre>
<p><code>max-age</code> talks to the browser; <code>s-maxage</code> talks to shared caches —the CDN— and takes precedence over it. That separation is what lets you ship changes fast: the user&#39;s browser holds nothing, the edge holds five minutes, and purging the edge is under your control. Purging someone&#39;s browser never is.</p>
<blockquote>
<p>💡 If your HTML leaves the origin with <code>max-age=3600</code>, you don&#39;t have a CDN problem: you have users stuck on an old version for an hour with no way to fix it. <code>max-age=0, s-maxage=3600</code> caches just as well and can actually be rolled back.</p>
</blockquote>
<h2>Invalidation is plan B; versioning is plan A</h2>
<p>Invalidations are the tool everyone reaches for first and the one you should need least. They&#39;re asynchronous, they take time, the first 1,000 paths a month are free and then billed, and above all: a <code>/*</code> on every deploy empties the whole cache and sends all your traffic to the origin at once, right at the moment you just deployed. That&#39;s the worst possible combination.</p>
<p>The alternative is making invalidation unnecessary. If your assets carry a content hash in the filename, a deploy doesn&#39;t modify files: it publishes new files, at new URLs, that nobody has cached. There&#39;s nothing to invalidate. The only thing that moves is the entry document, and that one already has a short <code>s-maxage</code>.</p>
<p>When it&#39;s still needed, I scope it:</p>
<pre><code class="language-bash"># Plan B, and scoped: entry documents only
aws cloudfront create-invalidation \
  --distribution-id E2XXXXXXXXXXXX \
  --paths &#39;/index.html&#39; &#39;/blog/*&#39;
</code></pre>
<h2><code>stale-while-revalidate</code>: let the origin stop suffering</h2>
<p>Without <code>stale-while-revalidate</code>, TTL expiry is a cliff: the request that arrives right after it waits for the origin to respond in full. Under real traffic, several requests arrive at once and all of them hit the origin for the same URL.</p>
<p>With SWR, the edge serves the stale copy immediately and revalidates in the background. The user never pays the refresh latency, and your p99 stops showing periodic spikes that correlate with nothing. Add <code>stale-if-error</code> and you get a degraded mode for free: if the origin returns a 5xx, the last good copy keeps being served instead of propagating the error.</p>
<p>If the origin still feels the revalidation traffic, Origin Shield adds an intermediate cache layer that consolidates those requests. It costs extra: I turn it on when the data asks for it, not by default.</p>
<h2>How I measure before touching anything</h2>
<p>The console&#39;s <code>CacheHitRate</code> tells you that you have a problem, not where it is. That&#39;s what the logs are for: every line carries <code>x-edge-result-type</code> with <code>Hit</code>, <code>RefreshHit</code>, <code>Miss</code>, <code>Error</code>. Grouping by URI and by that field gives you the diagnosis in one query.</p>
<pre><code class="language-sql">SELECT uri, x_edge_result_type, count(*) AS n
FROM cloudfront_logs
WHERE date &gt;= current_date - interval &#39;7&#39; day
GROUP BY 1, 2
ORDER BY n DESC
LIMIT 50;
</code></pre>
<p>If the same URI shows up near the top with <code>Miss</code> and high volume, there are only two explanations: the cache key is fragmenting it, or the origin is sending a <code>Cache-Control</code> that prevents caching. Both take ten minutes to fix once you know which one it is.</p>
<h2>What I&#39;d do today, in this order</h2>
<p>Decide <code>Cache-Control</code> at the origin per route type before anything else. Start the cache key empty and add only what&#39;s justifiable. Put content hashes in asset filenames so invalidation becomes exceptional. <code>stale-while-revalidate</code> on anything that&#39;s HTML. And measure with <code>x-edge-result-type</code> before pulling the next lever.</p>
<p>A badly configured CDN isn&#39;t neutral: it adds a hop, a layer for bugs to hide in, and a transfer bill you pay twice. Configured well, it&#39;s one of the few things in AWS where savings and latency improve in the same direction. The difference between those two versions isn&#39;t the product: it&#39;s three decisions that fit in an afternoon.</p>
]]></content:encoded>
    </item>
    <item>
      <title>INP: why your app feels slow even when it loads fast</title>
      <link>https://yohangel.com/en/blog/inp-responsividad-interacciones/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/inp-responsividad-interacciones/</guid>
      <description>A green LCP does not mean your app responds well. INP measures what the user feels on every click, and fixing it changes how your product is perceived.</description>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>Frontend</category>
      <content:encoded><![CDATA[<p><strong>Your LCP is green, the bundle is optimized, Lighthouse hands you a 95, and someone on the team still says &quot;the app feels slow.&quot; They&#39;re not wrong. They&#39;re describing something your load metric can&#39;t see: what happens when they tap a button and the interface takes a beat to react. That feeling has a name, and since 2024 it&#39;s a Core Web Vital: INP, Interaction to Next Paint. It measures the worst-case latency between a user interacting and the screen responding. A product can load blazingly fast and still feel like junk, because loading and responding are two different problems. This is the one I&#39;ve had the hardest time explaining to people who only stare at the load number.</strong></p>
<h2>What INP measures, and why LCP doesn&#39;t cover it</h2>
<p>LCP (Largest Contentful Paint) answers &quot;how long until the main content showed up?&quot; It&#39;s a startup problem: it happens once, at the beginning. INP is something else entirely. Across the whole session, every click, every tap, every keystroke fires a cycle of &quot;browser receives the event → runs your JavaScript → repaints.&quot; INP keeps the worst of those latencies, and that&#39;s your grade.</p>
<p>The distinction matters because modern apps live in interaction, not in load. A dashboard, an editor, a filter panel: the user loads once and then interacts a thousand times. If every filter they toggle freezes the UI for 300ms, your perfect LCP won&#39;t save you. A good INP threshold is 200ms; above 500ms it feels broken.</p>
<h2>The culprit is almost always the same: long tasks</h2>
<p>The browser runs your JavaScript on a single thread. While a function is running, it can&#39;t repaint or handle another event until that function returns. If clicking kicks off a computation or a render that takes 250ms, then for those 250ms the interface is dead: the button doesn&#39;t show its active state, nothing moves. That&#39;s a &quot;long task,&quot; and it&#39;s 90% of the INP problems I&#39;ve debugged.</p>
<p>The classic pattern: a search input that, on every keystroke, filters and re-renders a list of 2,000 items. Each press rebuilds the whole tree before letting the browser paint the very text you&#39;re typing. The result is an input that stutters even though the filtering is trivial.</p>
<pre><code class="language-js">// Find where the time goes: log any task that blocks &gt;50ms
new PerformanceObserver((list) =&gt; {
  for (const entry of list.getEntries()) {
    console.warn(&#39;Long task:&#39;, Math.round(entry.duration), &#39;ms&#39;, entry);
  }
}).observe({ type: &#39;longtask&#39;, buffered: true });
</code></pre>
<p>Before optimizing anything, I measure. The <code>longtask</code> <code>PerformanceObserver</code> tells you which interactions block the thread and for how long. Don&#39;t guess: it&#39;s almost never what you think.</p>
<h2>Fixing it in React: separate the urgent from what can wait</h2>
<p>The key idea is that not everything an interaction triggers is equally urgent. When you type into a search box, showing the character you just typed is extremely urgent; recomputing the filtered list can wait 50ms without anyone noticing. React 18 gave me the tools to say exactly that.</p>
<p><code>useDeferredValue</code> marks a value as &quot;low priority&quot;: React paints the urgent input (the text in the box) first and recomputes the derived work afterward, without blocking.</p>
<pre><code class="language-jsx">function Search({ items }) {
  const [query, setQuery] = useState(&#39;&#39;);
  const deferred = useDeferredValue(query);

  // Recomputed from the deferred value, not on every keystroke
  const filtered = useMemo(
    () =&gt; items.filter((i) =&gt; i.name.includes(deferred)),
    [items, deferred],
  );

  return (
    &lt;&gt;
      &lt;input value={query} onChange={(e) =&gt; setQuery(e.target.value)} /&gt;
      &lt;List items={filtered} /&gt;
    &lt;/&gt;
  );
}
</code></pre>
<p>The input responds instantly because its render doesn&#39;t wait for the filtering. For heavier actions — switching tabs, applying a filter that reshuffles half the screen — I reach for <code>useTransition</code>, which wraps the expensive update and lets the click show its feedback before the work starts:</p>
<pre><code class="language-jsx">const [pending, startTransition] = useTransition();

function onFilter(next) {
  setActiveFilter(next);            // urgent: the button marks itself now
  startTransition(() =&gt; {
    setResults(recompute(next));    // non-urgent: doesn&#39;t block the click
  });
}
</code></pre>
<h2>When there&#39;s no React involved: yield the thread</h2>
<p>Sometimes the heavy work isn&#39;t a render, it&#39;s a loop: processing a CSV, transforming a large JSON. There the fix is to slice the long task into chunks and hand control back to the browser between them, so it can handle clicks and repaint. The modern way is <code>scheduler.yield()</code>; the timeless fallback is <code>setTimeout(0)</code>.</p>
<pre><code class="language-js">async function process(rows) {
  for (let i = 0; i &lt; rows.length; i++) {
    work(rows[i]);
    if (i % 100 === 0) await yieldToMain(); // breathe every 100 rows
  }
}

const yieldToMain = () =&gt;
  &#39;scheduler&#39; in window &amp;&amp; &#39;yield&#39; in scheduler
    ? scheduler.yield()
    : new Promise((r) =&gt; setTimeout(r, 0));
</code></pre>
<p>If the work is genuinely CPU-bound and doesn&#39;t need the DOM, the right answer isn&#39;t chunking but moving it into a Web Worker: another thread, zero blocking of the main one. But for most cases, yielding the thread every N iterations already turns a freeze into something imperceptible.</p>
<h2>What I learned by measuring, not guessing</h2>
<p>INP taught me to stop conflating &quot;fast to load&quot; with &quot;fast to use.&quot; They&#39;re different axes, and the user feels the second one most, because they spend 99% of their time interacting, not loading. The good news is it almost always gets fixed without rewriting anything: you find the long task with <code>longtask</code>, decide what&#39;s urgent and what can wait, and use the right tool — <code>useDeferredValue</code>, <code>useTransition</code>, yielding the thread, or a worker. My rule is simple: if an interaction blocks the thread for more than 200ms, it&#39;s not an optional micro-optimization, it&#39;s a quality-perception bug. And perception bugs are the ones that make people say &quot;I don&#39;t know, it feels slow&quot; and not come back.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Cheaper embeddings: trimming dimensions and quantizing without losing recall</title>
      <link>https://yohangel.com/en/blog/embeddings-cuantizacion-dimensiones/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/embeddings-cuantizacion-dimensiones/</guid>
      <description>How I decide how many dimensions to keep and at what precision, and why I almost always end up with a two-phase scheme: small vectors on top, precise ones underneath.</description>
      <pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>IA</category>
      <content:encoded><![CDATA[<p><strong>A vector index gets expensive much sooner than people expect. Not because of the embedding model, which already costs pennies, but because of memory: if you store 3072-dimensional vectors in float32, each one takes 12 KB, and a million vectors is 12 GB of raw data alone, before you count the index graph. What I learned building the matching engine at JXBS is that this size is negotiable along two independent axes —how many dimensions you store and at what precision you store each one— and that cutting on both degrades recall far less than intuition suggests, as long as you use the small vector to search and the big one to decide.</strong></p>
<h2>The two axes: dimensions and precision</h2>
<p>When you want an embedding to take up less space, there are exactly two levers. The first is truncating dimensions: keeping the first N components of the vector instead of all of them. The second is quantizing: lowering the precision of each component, from <code>float32</code> to <code>int8</code> or even to a single bit.</p>
<p>These levers multiply. A 1024-dimensional vector in int8 takes 1 KB versus 4 KB for the same vector in float32, and versus 12 KB for the full 3072-dimensional float32 vector. Twelve times less memory. The question isn&#39;t whether that saves money —obviously it does— but how much recall you lose along the way.</p>
<h2>Truncating dimensions isn&#39;t as barbaric as it sounds</h2>
<p>Truncating an arbitrary vector destroys information unpredictably: there&#39;s no guarantee the first components are the important ones. What changed the landscape is <em>matryoshka</em>-style training, where the model is explicitly trained so that prefixes of the vector are, on their own, useful representations. The first dimensions carry the coarse signal; the last ones refine nuance.</p>
<p>If your model supports this scheme —OpenAI&#39;s main models and several open ones do— truncating is literally slicing the array and renormalizing:</p>
<pre><code class="language-typescript">function truncate(vector: number[], dims: number): number[] {
  const slice = vector.slice(0, dims);
  const norm = Math.hypot(...slice);
  return slice.map((v) =&gt; v / norm); // renormalizing is mandatory
}
</code></pre>
<p>That <code>renormalize</code> is not optional. If you use cosine similarity with normalized vectors and truncate without normalizing again, magnitudes stop being comparable and distances get dirty. It&#39;s the most common error and the quietest one: nothing blows up, results just get worse.</p>
<p>If your model is <strong>not</strong> trained with matryoshka, don&#39;t truncate. There the honest reduction goes through PCA or a learned projector, and that&#39;s one more component to maintain, version, and reindex.</p>
<h2>Quantizing: int8 nearly free, binary with caveats</h2>
<p>Dropping from float32 to int8 preserves surprisingly high recall —in my tests the loss has been marginal— in exchange for a quarter of the memory. The trick is calibration: you need to know the actual range of your components in order to map it to int8&#39;s range, and that range is computed over a representative sample of your corpus, not over the first batch you happen to have on hand.</p>
<p>Binary quantization is another story. You reduce each component to one bit based on its sign, take up 32 times less space, and compare with Hamming distance, which is an XOR and a popcount: brutally fast. But you lose recall visibly. Nobody serious uses binary as the final answer; it&#39;s used as a filter.</p>
<blockquote>
<p>💡 The right question isn&#39;t &quot;what precision do I use?&quot; but &quot;what precision do I use <em>in each phase</em>?&quot;. Searching and deciding are different problems and deserve different representations.</p>
</blockquote>
<h2>The pattern I use: search cheap, decide expensive</h2>
<p>Instead of choosing a single format, I store two. A small, quantized vector to traverse the index, and the full float32 vector —or something close— to rerank the few candidates that survive.</p>
<p>The first phase retrieves, say, the 200 nearest candidates using the cheap vector. It&#39;s a <em>recall</em> phase: it only has to guarantee the good ones are inside those 200, not that they&#39;re well ordered. The second phase takes those 200, computes similarity with the full-precision vector and orders them for real. That second phase touches 200 vectors, not a million, so index memory doesn&#39;t constrain it: the full vectors can live on disk or in a regular Postgres table.</p>
<pre><code class="language-sql">-- Phase 1: cheap recall over the quantized vector (indexed)
WITH candidates AS (
  SELECT id
  FROM documents
  ORDER BY embedding_int8 &lt;=&gt; $1::vector
  LIMIT 200
)
-- Phase 2: rerank the 200 with the full-precision vector
SELECT d.id, d.title, d.embedding_full &lt;=&gt; $2::vector AS distance
FROM documents d
JOIN candidates c ON c.id = d.id
ORDER BY distance
LIMIT 10;
</code></pre>
<p>This is exactly the same idea as putting a reranker on top of semantic search, but one floor down: instead of an expensive model reordering, it&#39;s an expensive vector reordering. And you can combine both.</p>
<h2>How I decide the cutoff</h2>
<p>There&#39;s no magic number. What I do is measure. I freeze a set of queries with their known relevant results, define recall@10 with the full float32 vector as the reference, and try configurations: 3072/float32, 1024/float32, 1024/int8, 512/int8, binary+rerank. Each one gives a (recall, memory) pair. With that table in front of you the decision stops being an argument about opinions and becomes an explicit trade-off choice.</p>
<p>My default bias, when the corpus is large: medium dimensions with int8 in the index, full vector stored separately for the second phase. I chose that because the index&#39;s memory cost drops nearly an order of magnitude, in exchange for a two-step query instead of one and for keeping two columns in sync when you reindex. That&#39;s the real price, and it needs saying: <strong>every compression scheme is a pending migration</strong>. The day you change embedding models, you reindex both columns.</p>
<p>And a final warning: don&#39;t optimize this before you have the problem. With a hundred thousand vectors, full float32 fits in memory without breaking a sweat and you won&#39;t notice the difference. Compression is an answer to a concrete constraint, not a virtue in itself.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Forms that work without JavaScript (and better with it)</title>
      <link>https://yohangel.com/en/blog/formularios-progressive-enhancement/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/formularios-progressive-enhancement/</guid>
      <description>The progressive enhancement pattern I use for forms: the platform solves the base case, JavaScript only adds layers, and neither path is second-class.</description>
      <pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>Frontend</category>
      <content:encoded><![CDATA[<p><strong>A form is the piece where the difference shows most between a site built on top of the platform and one built against it. The browser already knows how to submit data, show validation errors, manage focus state, and navigate to the result. It has known this for decades. And yet almost everyone&#39;s default reflex, mine included for years, is to intercept the <code>submit</code>, <code>preventDefault()</code>, wire up a React state, and reimplement all of that worse. This article is about the opposite pattern: writing the form so it works without a single line of JavaScript, and then using JavaScript to make it better —not to make it possible.</strong></p>
<h2>The base case: a <code>&lt;form&gt;</code> that already works</h2>
<p>A form with <code>action</code> and <code>method</code> sends data to the server and navigates to the result. No JS. Inputs with <code>required</code>, <code>type=&quot;email&quot;</code>, <code>minlength</code> or <code>pattern</code> validate on the client without a line of code. The associated <code>&lt;label&gt;</code> gives you accessibility for free. The submit button shows a native browser loading state.</p>
<pre><code class="language-html">&lt;form action=&quot;/api/contact&quot; method=&quot;post&quot;&gt;
  &lt;label for=&quot;email&quot;&gt;Email&lt;/label&gt;
  &lt;input id=&quot;email&quot; name=&quot;email&quot; type=&quot;email&quot; required /&gt;
  &lt;button type=&quot;submit&quot;&gt;Send&lt;/button&gt;
&lt;/form&gt;
</code></pre>
<p>That works on a phone with half a network, while your bundle is still downloading, with JS disabled, or when a third-party script blows up and takes hydration down with it. On an Astro site, where most islands don&#39;t even hydrate, this isn&#39;t a hypothetical: it&#39;s the page&#39;s normal behavior.</p>
<h2>Validation can&#39;t live in one place only</h2>
<p>This is where discipline usually breaks. People validate on the client with a library, and on the server they validate again with different hand-written logic. Two sources of truth, which diverge the moment someone changes a rule.</p>
<p>What I do is have a single schema —Zod— and use it on both sides. On the server it&#39;s the truth; the client is a courtesy to give fast feedback. And crucially: if the JavaScript didn&#39;t load, the server still validates and returns a page with the errors. The base case is never left unprotected.</p>
<pre><code class="language-typescript">// shared
export const contact = z.object({
  email: z.string().email(&#39;Invalid email&#39;),
  message: z.string().min(10, &#39;Tell me a bit more&#39;),
});

// server: the only truth
export async function POST({ request }: { request: Request }) {
  const data = Object.fromEntries(await request.formData());
  const res = contact.safeParse(data);
  if (!res.success) {
    // no JS: render the page back with the errors
    return renderWithErrors(data, res.error.flatten().fieldErrors);
  }
  await save(res.data);
  return new Response(null, { status: 303, headers: { Location: &#39;/thanks&#39; } });
}
</code></pre>
<p>Note the <code>303</code> with <code>Location</code>. That redirect after a successful POST is the POST/Redirect/GET pattern, and it prevents refreshing the page from resubmitting the form. It&#39;s from 1998 and it&#39;s still correct.</p>
<blockquote>
<p>💡 If your form breaks when JavaScript fails, you didn&#39;t have a form. You had a widget that looked like one.</p>
</blockquote>
<h2>Adding JavaScript on top, not underneath</h2>
<p>With a solid base case, JS comes in to improve things. And it does so by listening to the <code>submit</code> of the form that <em>already exists</em>, not by building a new one.</p>
<pre><code class="language-typescript">form.addEventListener(&#39;submit&#39;, async (e) =&gt; {
  e.preventDefault(); // only runs if this script loaded
  const data = new FormData(form);
  const res = contact.safeParse(Object.fromEntries(data));
  if (!res.success) return paintErrors(res.error.flatten().fieldErrors);

  button.disabled = true;
  const r = await fetch(form.action, { method: &#39;POST&#39;, body: data });
  button.disabled = false;
  if (r.redirected) location.assign(r.url);
});
</code></pre>
<p>The key is the mental ordering: <code>preventDefault()</code> is an <em>optional optimization</em> that only happens if the script ran. If it doesn&#39;t run, the browser does its usual job. There&#39;s no broken path, there&#39;s a less polished path.</p>
<p>And notice something else: I use <code>new FormData(form)</code> instead of reading React state. The DOM is already the form&#39;s state. Keeping a copy in <code>useState</code> synced on every <code>onChange</code> is work the platform does for free, and it&#39;s the cause of half the form bugs I&#39;ve debugged: controlled inputs that lose the cursor, browser autofill that doesn&#39;t fire the event, password managers that fill fields and React state never finds out.</p>
<h2>What I gain and what I pay</h2>
<p>I gain real robustness: the form works before hydration, on bad networks, with JS down. I gain accessibility nearly for free, because native elements already come with their semantics. I gain less code: no form state reducer, no form library, no duplicated validation.</p>
<p>What I pay is that certain very rich interactions cost more. A field with remote autocomplete, drag-and-drop file uploads, or a multi-step wizard with complex state don&#39;t come out of a bare <code>&lt;form&gt;</code>. There I do build a controlled component, but in a localized way: the complex island is controlled, and the rest of the form is still the platform. I chose this because complexity stays contained where I need it, in exchange for having two field-handling styles coexisting in the same file, which is ugly if you don&#39;t document it.</p>
<p>I also pay that the &quot;no JS&quot; case has to be <em>tested</em>. If you don&#39;t test it, it rots. On projects where this matters to me, I have a test that submits the form with a plain HTTP request, no browser, and checks that the response comes back with the errors rendered. It&#39;s a cheap test and it&#39;s the only one that guarantees the base case is still alive.</p>
<p>The conclusion isn&#39;t &quot;don&#39;t use JavaScript&quot;. It&#39;s that JavaScript should be the second layer, not the foundation. A form that only exists once the bundle loads is a form with a dependency you never agreed to take on.</p>
]]></content:encoded>
    </item>
    <item>
      <title>What's going to happen to software development: predictions with dates (July 2026)</title>
      <link>https://yohangel.com/en/blog/futuro-desarrollo-predicciones-2026/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/futuro-desarrollo-predicciones-2026/</guid>
      <description>What to expect in the coming months, in 2027 and toward 2030: agents, the future of web and apps, and what happens to programmers. With data, not vibes.</description>
      <pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>IA</category>
      <content:encoded><![CDATA[<p><strong>Predicting the future of software has become an extreme sport: the best minds in the industry have been failing at it in public for three years. So before placing my own bets, let&#39;s audit theirs — because the mistakes of the people who know the most contain the best clues about what&#39;s coming.</strong></p>
<h2>First: the prophets&#39; track record</h2>
<p>Let&#39;s review famous predictions and their actual status as of July 2026:</p>
<table>
<thead>
<tr>
<th>Who and when</th>
<th>Prediction</th>
<th>Status today</th>
</tr>
</thead>
<tbody><tr>
<td>Dario Amodei, Mar 2025</td>
<td>&quot;AI will write 90% of code in 3–6 months&quot;</td>
<td>True inside Anthropic (~90–100%); false industry-wide on that timeline</td>
</tr>
<tr>
<td>Zuckerberg, Jan 2025</td>
<td>&quot;In 2025 we&#39;ll have AI working as a mid-level engineer&quot;</td>
<td>Not as an autonomous employee; Meta cut 8,000 jobs and moved 7,000 into AI</td>
</tr>
<tr>
<td>Jensen Huang, Feb 2024</td>
<td>&quot;Don&#39;t learn to code&quot;</td>
<td>Demand for AI-skilled devs hasn&#39;t stopped growing</td>
</tr>
<tr>
<td>Sam Altman, Jan 2025</td>
<td>&quot;Agents will join the workforce in 2025&quot;</td>
<td>Not as &quot;employees&quot;; yes as omnipresent tools</td>
</tr>
<tr>
<td>Sundar Pichai, 2024–26</td>
<td>25% → 50% → 75% of Google&#39;s new code by AI</td>
<td>Came true, on schedule</td>
</tr>
</tbody></table>
<p>The pattern is crystal clear, and it&#39;s the key to this whole post: <strong>the optimists get the direction right and the timeline wrong — systematically too fast on the social side and too slow on the technical side.</strong> Nobody predicted agents would write code this well this soon; everybody overestimated how fast organizations would absorb it.</p>
<h2>The number that anchors any serious prediction</h2>
<p>If you can only watch one metric, watch METR&#39;s: the length of tasks an agent can complete autonomously (at a 50% success rate) doubles every few months — 196 days on the full historical average, but accelerating toward ~3 months in the post-2024 data. In January 2026, the best models completed tasks of ~5 hours of human work; unofficial trackers already place current models above 14 hours.</p>
<p>If the trend holds — and it has held for six years — by mid-2027 we&#39;re talking about agents completing tasks worth a week of human work. That single curve explains almost everything that follows.</p>
<h2>The coming months (rest of 2026): agentic consolidation</h2>
<p>What I consider practically certain between now and December, because it&#39;s already underway:</p>
<ul>
<li><strong>Orchestration becomes the job.</strong> Karpathy named it &quot;agentic engineering&quot; and he was right: the conversation is no longer &quot;which editor do you use&quot; but &quot;how many agents do you direct in parallel.&quot; Fleet tooling (dynamic workflows, remote tasks, cloud agents) becomes standard.</li>
<li><strong>Bill shock.</strong> This year&#39;s billing changes (Copilot credits, usage limits everywhere) are the symptom: agentic compute cost becomes a serious line item in every team&#39;s budget. FinOps tooling for agents will emerge.</li>
<li><strong>MCP locks in as the universal standard.</strong> The July spec (embedded apps, long-running tasks, serious OAuth) plus Linux Foundation governance make it the TCP/IP of agents. If your product doesn&#39;t expose MCP in 2027, it doesn&#39;t exist for agents.</li>
<li><strong>Verification becomes the official bottleneck.</strong> With benchmarks like SWE-bench saturated (95% for the best model), the problem is no longer generating code — it&#39;s reviewing it. 66% of devs say they lose more time fixing &quot;almost right&quot; AI code than writing. That&#39;s where the next wave of tools is.</li>
</ul>
<h2>2027: the year of the interface</h2>
<p>My central bet for 2027 isn&#39;t about code — it&#39;s about interfaces. The pieces are already on the table: agentic browsers (Atlas, Comet, Dia) with tens of millions of users, ChatGPT as an app platform with built-in checkout, and WebMCP so sites can expose actions directly to agents.</p>
<p>What that means in practice:</p>
<ul>
<li><strong>Your website will have two audiences: humans and agents.</strong> Just as mobile-first happened, agent-first is coming: clean APIs, explicit semantics, actions exposed via MCP/WebMCP. A growing share of your &quot;visits&quot; will never see your CSS.</li>
<li><strong>Generative UI will find its place — which is not everywhere.</strong> The dream of &quot;the app is generated on the fly and deleted after&quot; will collide with what Nielsen has warned about for years: an interface that changes every time is an interface no one can learn. The equilibrium: your UI as a reference implementation, with data and actions open so the user&#39;s agent can compose its own.</li>
<li><strong>Agentic commerce becomes normal.</strong> With hundreds of millions of weekly assistant users and standardized purchase protocols, &quot;I asked my agent to order it&quot; stops sounding weird. SEO mutates: from ranking pages to ranking actions and data.</li>
</ul>
<blockquote>
<p>💡 If you maintain a web product, the question for 2027 isn&#39;t &quot;do I have a mobile app?&quot; but &quot;can an agent use my product without a screen?&quot;. Whoever has a good API and good semantics wins the new channel for free.</p>
</blockquote>
<h2>2027–2028: what happens to programmers</h2>
<p>This is where the noise is loudest and where the data says something more nuanced than the headlines:</p>
<ul>
<li><strong>The pyramid restructures; it doesn&#39;t disappear.</strong> Software engineers are now 55% of Big Tech hiring — more than in 2019 (46%). But new grads are only 7% of those hires, half the pre-pandemic share. More senior hiring, less junior: the pyramid is inverting.</li>
<li><strong>The jobs moved; they didn&#39;t vanish.</strong> US dev postings are up 14% year over year, but 71% of that increase is senior roles and 37% carry &quot;AI&quot; in the title. Average AI engineer compensation sits around $242K, and agent-focused roles are growing +136% year over year.</li>
<li><strong>The 2029 shortage is being manufactured today.</strong> CS enrollment falling 8% a year, bootcamps closing in waves, companies not hiring juniors. Nobody is training the seniors of five years from now. Concrete prediction: around 2028–2029, a senior-talent-shortage panic and &quot;AI apprenticeship&quot; programs everywhere, hiring juniors again — with a different profile: systems design and verification, not syntax.</li>
<li><strong>The &quot;because of AI&quot; layoffs will continue — and won&#39;t be only because of AI.</strong> 2026 has seen ~120,000 tech layoffs with AI as a partial excuse — a mix of real automation, past overhiring and margin pressure. Untangling how much is which will be impossible, and the &quot;AI took my job&quot; narrative will coexist with record demand for engineers who know how to direct it.</li>
</ul>
<h2>2028–2030: the three scenarios</h2>
<p>More than two years out, the honest move is scenarios with probabilities, not certainties. Mine:</p>
<ul>
<li><strong>Continued acceleration (~35%).</strong> The METR curve holds with no ceiling: agents with weeks of autonomy in 2027, a functional &quot;digital employee&quot; toward 2028–29. This is the labs&#39; scenario (though even Amodei and Altman have softened the apocalypse talk this year — curiously, on their way to IPOs).</li>
<li><strong>Useful plateau (~50%).</strong> The one I find most likely. Code generation keeps improving but the hard problems — persistent memory, continual learning, long-horizon coherence — prove stubborn, as Marcus and LeCun warn (LeCun left Meta to found a world-models startup precisely over this). Agents become infrastructure the way the cloud did: transformative, not apocalyptic. &quot;AI 2027&quot;-style forecasts have already slipped to &quot;early 2030s,&quot; and Metaculus puts 50% on AGI by 2033.</li>
<li><strong>Partial winter (~15%).</strong> Compute spending doesn&#39;t find returns at the promised pace, a hard valuation correction, consolidation. Note: even here, the capabilities already deployed don&#39;t go away — nobody is going back to writing CRUD by hand.</li>
</ul>
<h2>My concrete predictions, so I can fail in public</h2>
<p>Since I&#39;ve spent this post auditing other people&#39;s predictions, it&#39;s only fair to leave mine in writing, dated, so they can be audited:</p>
<pre><code class="language-text">// Predictions — July 2026 (review in July 2027)
// 1. By Dec 2026: agents reliably complete tasks worth
//    ~2 days of human work.                          [80%]
// 2. By Jun 2027: &gt;50% of new code at tech companies is
//    AI-generated (industry average, not just labs).  [70%]
// 3. By 2027: at least one top-100 app drops its
//    traditional UI for an agentic/conversational one. [60%]
// 4. By 2028: junior hiring recovers in an AI-apprentice
//    format after the shortage panic.                 [65%]
// 5. By 2030: MORE people work in software than in 2025
//    (counting the new roles).                        [75%]
// 6. AGI / drop-in digital worker before 2030.        [25%]
</code></pre>
<h2>The honest closing</h2>
<p>Prediction number 5 is the one that really matters, and the one I&#39;m most convinced will come true. Every time the cost of creating software has collapsed — compilers, open source, the cloud — the world didn&#39;t want less software: it wanted orders of magnitude more. Software demand has always been supply-constrained, and we just made the supply nearly infinite.</p>
<p>What disappears isn&#39;t the programmer; it&#39;s the programmer defined as &quot;a person who translates specifications into syntax.&quot; What&#39;s being born looks more like a systems director who decides, verifies and answers for the result. The next three years will be uncomfortable, uneven and full of exaggerated headlines in both directions. But if I had to pick one moment in history to know how to build software, I&#39;d pick exactly this one.</p>
]]></content:encoded>
    </item>
    <item>
      <title>The NAT Gateway nobody asked for: your VPC's silent bill</title>
      <link>https://yohangel.com/en/blog/nat-gateway-vpc-endpoints-costes/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/nat-gateway-vpc-endpoints-costes/</guid>
      <description>Why the most expensive component in many AWS accounts is a networking piece nobody remembers designing, and how VPC endpoints change that equation.</description>
      <pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>AWS</category>
      <content:encoded><![CDATA[<p><strong>In almost every AWS account I&#39;ve reviewed there&#39;s a line on the bill that surprises everyone: the NAT Gateway. It&#39;s not a service anyone consciously chose. It shows up because the VPC creation wizard puts it there by default, because your Fargate tasks live in private subnets and need to reach the internet, and because nobody looked afterward. It charges per hour <em>and</em> per gigabyte processed, in both directions. And the worst part: much of that traffic isn&#39;t even going to the internet —it&#39;s going to S3, to DynamoDB, to Secrets Manager— and you could have pulled it out of there with a component that costs zero.</strong></p>
<h2>Why the NAT Gateway exists</h2>
<p>A private subnet, by definition, has no route to an Internet Gateway. That&#39;s what makes it private: nothing outside can initiate a connection inward. But your containers do need to get out: to pull an image from ECR, to read a secret, to call a third-party API.</p>
<p>The NAT Gateway solves that: it lives in a public subnet, private subnets route their outbound traffic to it, and it translates addresses so responses come back. It works, it&#39;s managed, it&#39;s highly available within its AZ. And that&#39;s why everyone puts it there and forgets about it.</p>
<p>The problem is its pricing model. You pay an hourly rate per NAT Gateway —and since each one lives in an AZ, the recommended high-availability practice asks for one per AZ, so multiply— plus a rate for every gigabyte that crosses it. That per-gigabyte charge applies to <em>processed</em> traffic, and it accumulates without anyone watching.</p>
<h2>The uncomfortable discovery: much of that traffic is internal</h2>
<p>Here&#39;s the detail that changes the analysis. When your Fargate task in a private subnet writes an object to S3, that traffic exits through the NAT Gateway, goes to S3&#39;s public endpoint, and comes back. You&#39;re paying transfer to talk to an AWS service that sits in the same region as you.</p>
<p>Same with DynamoDB, with SQS, with Secrets Manager, with CloudWatch Logs —and CloudWatch Logs is the big silent suspect, because if you have verbose logs, every line is billed traffic— and with ECR every time a task starts and pulls image layers.</p>
<p>In well-populated serverless or container architectures, most outbound traffic isn&#39;t really going to the internet. It&#39;s going to AWS.</p>
<h2>VPC endpoints: two flavors, very different prices</h2>
<p>A VPC endpoint lets your private subnet talk to an AWS service without going through the internet or the NAT. There are two types, and confusing them is expensive.</p>
<p><strong>Gateway endpoints</strong> exist only for S3 and DynamoDB. They&#39;re an entry in the route table. <strong>They cost nothing</strong>: not per hour, not per gigabyte. There is no defensible reason not to have them if you use S3 or DynamoDB from private subnets. It&#39;s money thrown away.</p>
<p><strong>Interface endpoints</strong> (PrivateLink) are for everything else: SQS, Secrets Manager, ECR, CloudWatch, KMS. They create a network interface in your subnet with a private IP. These do cost: an hourly rate per endpoint per AZ, plus a per-gigabyte charge —quite a bit cheaper than the NAT&#39;s, but not zero.</p>
<blockquote>
<p>💡 S3 and DynamoDB gateway endpoints are free. If your private subnets don&#39;t have them, you&#39;re paying NAT Gateway to talk to S3. That&#39;s the first place to look, and usually the most profitable.</p>
</blockquote>
<h2>How I decide which endpoints to create</h2>
<p>With interface endpoints the math isn&#39;t automatic, because they have a fixed cost per hour per AZ. An endpoint pays off when the savings in NAT traffic exceed that fixed rate. The way to know is to look at the data, not guess: I enable Flow Logs on the NAT Gateway&#39;s elastic network interface and group by destination IP, resolving those IPs against AWS&#39;s published ranges to know which service is which.</p>
<p>With that table in front of you the decision is arithmetic. If 60% of your NAT traffic is to S3, a free gateway endpoint takes all of it. If another 20% is CloudWatch Logs, the interface endpoint pays for itself. If there&#39;s 5% to KMS, the fixed cost probably isn&#39;t worth it.</p>
<p>In OpenTofu it&#39;s so little code it&#39;s almost embarrassing not to have it:</p>
<pre><code class="language-hcl"># Free. No excuses.
resource &quot;aws_vpc_endpoint&quot; &quot;s3&quot; {
  vpc_id            = aws_vpc.main.id
  service_name      = &quot;com.amazonaws.${var.region}.s3&quot;
  vpc_endpoint_type = &quot;Gateway&quot;
  route_table_ids   = aws_route_table.private[*].id
}

# Paid: creates one ENI per subnet. Justify it with data.
resource &quot;aws_vpc_endpoint&quot; &quot;logs&quot; {
  vpc_id              = aws_vpc.main.id
  service_name        = &quot;com.amazonaws.${var.region}.logs&quot;
  vpc_endpoint_type   = &quot;Interface&quot;
  subnet_ids          = aws_subnet.private[*].id
  security_group_ids  = [aws_security_group.endpoints.id]
  private_dns_enabled = true # without this, your SDK keeps going to the public endpoint
}
</code></pre>
<p>That <code>private_dns_enabled</code> is what makes it work without touching code: it resolves the service&#39;s public name to the endpoint&#39;s private IP. If you leave it at <code>false</code>, you create the endpoint, pay for it, and your application keeps happily exiting through the NAT. It&#39;s an error you only see on the bill.</p>
<h2>The most uncomfortable question: do you need the NAT?</h2>
<p>Before optimizing the NAT Gateway it&#39;s worth asking whether it&#39;s needed at all. If your containers only talk to AWS services and never call a third-party API, you can cover everything with endpoints and remove it. It&#39;s the only optimization that takes that cost to zero.</p>
<p>In practice there&#39;s almost always <em>something</em> going out to the internet: a webhook, a payment gateway, an LLM provider. So the NAT stays, but with most traffic diverted, and its bill goes from being a mystery to being a small, explainable line.</p>
<p>The honest trade-off: you add networking components that have to be maintained, versioned in your IaC, and debugged when a misconfigured security group makes an endpoint throw timeouts instead of clear errors. In exchange, internal traffic stops touching the internet —which is also better security posture— and the bill stops growing with your log volume. For me it&#39;s always been worth it. But decide it with Flow Logs, not with this article.</p>
]]></content:encoded>
    </item>
    <item>
      <title>How to survive as a programmer in the AI era (without turning cynical or naive)</title>
      <link>https://yohangel.com/en/blog/sobrevivir-programador-era-ia/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/sobrevivir-programador-era-ia/</guid>
      <description>AI isn't coming for your job — it's coming for your tasks. What's really changing, which skills are appreciating, and which ones are quietly losing value.</description>
      <pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>IA</category>
      <content:encoded><![CDATA[<p><strong>I&#39;ve been writing software for years and I&#39;ve never seen the profession this divided: half swear there won&#39;t be a single programmer left in two years, the other half swear it&#39;s all hype. Both camps are wrong — and while they argue, the ground is shifting under everyone.</strong></p>
<p>This isn&#39;t a motivational post or a list of &quot;5 courses so you don&#39;t fall behind.&quot; It&#39;s what I see from the inside — working every day with agents that write most of my code — about what&#39;s really changing, what&#39;s losing value, what&#39;s gaining it, and what I would do today depending on where you are in your career.</p>
<h2>First: separate signal from noise</h2>
<p>The confusion comes from mixing up two different questions. &quot;Can AI write code?&quot; — yes, it has for a while now, and better than most of us on well-scoped tasks. &quot;Can AI do an engineer&#39;s job?&quot; — no, because the job was never writing code.</p>
<p>The job was always something else: understanding an ambiguous problem, deciding what to build (and what not to), negotiating constraints, noticing that the ticket asks for one thing but the business needs another, and answering for what ships. None of that has been automated. What got automated was the part in the middle: translating decisions into syntax.</p>
<blockquote>
<p>💡 The sentence that keeps my head straight: AI isn&#39;t coming for your job, it&#39;s coming for your tasks. If your job was the sum of those tasks, you do have a problem. If your job was the judgment that ordered them, you just got multiplied.</p>
</blockquote>
<h2>What&#39;s losing value (and it hurts to say)</h2>
<p>Let&#39;s be honest about what&#39;s already worth less in the market:</p>
<ul>
<li><strong>Writing correct, clean code by hand.</strong> It was our identity. Today it&#39;s the baseline any well-directed agent produces.</li>
<li><strong>Knowing a framework by heart.</strong> Encyclopedic API knowledge used to be a competitive edge; now it&#39;s one prompt away from anyone.</li>
<li><strong>The junior who only executes tickets.</strong> This is the hardest hit and the most real one: &quot;take this well-defined task and bring it back done&quot; is exactly what agents do best.</li>
<li><strong>Typing speed as a productivity metric.</strong> Producing more lines no longer sets anyone apart. What does is producing fewer lines that survive longer.</li>
</ul>
<h2>What&#39;s appreciating</h2>
<p>The opposite list is more interesting, because it&#39;s where you should invest:</p>
<ul>
<li><strong>Review judgment.</strong> AI-generated code looks impeccable — formatted, sensibly named, commented. The danger lives in what <em>looks</em> correct. Reading code with suspicion is now worth more than writing it.</li>
<li><strong>Architecture and decomposition.</strong> An agent performs in direct proportion to how well-scoped the problem is. Slicing a system into pieces an AI can execute without breaking anything is the new senior skill.</li>
<li><strong>Business context.</strong> Understanding why something is being built is the one thing AI can&#39;t infer from your repo. The engineer who talks to product and to customers becomes impossible to replace.</li>
<li><strong>Verification.</strong> Tests, evals, observability, staging environments that resemble production. When code volume multiplies, the bottleneck moves to trust: how do I know this works?</li>
<li><strong>Accountability.</strong> Someone has to sign off on the deploy. The AI doesn&#39;t attend the postmortem. That &quot;someone&quot; is trading up.</li>
</ul>
<h2>The new workflow (mine, at least)</h2>
<p>My day-to-day no longer resembles what it was three years ago. The pattern that works for me has three phases, and none of them is &quot;writing code&quot;:</p>
<pre><code class="language-text">// My division of labor with agents
// 1. BEFORE — invest in context: clear specs, documented
//             conventions, explicit acceptance criteria.
// 2. DURING — steer, don&#39;t dictate: the agent proposes, I cut
//             scope, correct course, ask for alternatives.
// 3. AFTER  — verify with hostility: read the diff as if it were
//             written by someone trying to fool me, run the
//             tests, poke at the edges.
</code></pre>
<p>The ratio surprises anyone who hasn&#39;t lived it: I spend roughly 40% of my time in phase 1. The quality of what comes out of an agent is a nearly linear function of the quality of the context that goes in. Teams that adopt AI and see no improvement almost always fail there: they delegate the writing but don&#39;t invest in the specification.</p>
<h2>If you&#39;re just starting out: the uncomfortable advice</h2>
<p>The entry-level rung broke, and denying it helps no one. But notice the nuance: the rung broke, not the ladder. Companies still need seniors, and seniors aren&#39;t born — they&#39;re made. Sooner or later the market will have to rebuild the pipeline, and the juniors who survive this transition will be the ones who arrived differently:</p>
<ul>
<li><strong>Use AI to learn, not to avoid learning.</strong> Ask it to explain every line it generates. The difference between &quot;it works and I don&#39;t know why&quot; and &quot;it works and I know why&quot; is your entire career.</li>
<li><strong>Build complete things.</strong> A deployed side project, with real users and real bugs, teaches what no tutorial can: the part of the craft AI doesn&#39;t cover.</li>
<li><strong>Learn to actually debug.</strong> When the agent gets stuck — and it does — the person who can drop down to the log, the breakpoint and the protocol is the one who unblocks the team.</li>
<li><strong>Fundamentals don&#39;t expire.</strong> Networking, operating systems, databases, complexity. The frameworks AI has mastered change every year; the things they stand on don&#39;t.</li>
</ul>
<h2>If you&#39;ve been at this for years: your risk is different</h2>
<p>The senior doesn&#39;t compete against AI; they compete against the senior who uses it well. And there I see two symmetric mistakes. The first is rejection: &quot;I review better than any model&quot; — true today, irrelevant within two improvement cycles. The second is surrender: accepting everything the agent generates without reading it, until a production incident reminds you whose name was on the deploy.</p>
<p>The middle ground has a boring name: management. You direct a small fleet of agents the same way you used to coordinate people — with specs, with review, with standards. If you ever considered moving into management but didn&#39;t want to stop touching code, the good news is that this hybrid role was just invented and nobody has ten years of experience in it.</p>
<h2>What doesn&#39;t change</h2>
<p>After all the vertigo, one thing calms me down: software is still the discipline of deciding what should happen and making sure it does. The tools for getting there have changed more in three years than in the previous twenty, but the nature of the craft — judgment, accountability, translating between what people need and what the machine does — remains intact.</p>
<p>Surviving this era isn&#39;t about outrunning the AI. It&#39;s about moving one level above it — which is exactly the same move made by those who went from assembly to compilers, and from physical servers to the cloud. The ones who clung to the layer being automated had a rough time. The ones who moved up a layer couldn&#39;t keep up with all the work.</p>
]]></content:encoded>
    </item>
    <item>
      <title>The startup stack in July 2026: what I would use today to build a SaaS</title>
      <link>https://yohangel.com/en/blog/stack-startups-julio-2026/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/stack-startups-julio-2026/</guid>
      <description>An honest map of the ecosystem: coding agents, models, frameworks, infra and the AI layer. What I would pick today and why, with real prices.</description>
      <pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>IA</category>
      <content:encoded><![CDATA[<p><strong>If I had to build a SaaS from scratch today, the hard part wouldn&#39;t be building it — it would be choosing what to build it with. The ecosystem moves so fast that any comparison from six months ago is already wrong. This is my map as of July 2026: what I would use, what I would avoid, and where the pricing traps are.</strong></p>
<p>Fair warning: there is no perfect stack, only stacks that fit your team and your stage. What follows is biased toward what a small startup optimizes for: iteration speed, low maintenance, and predictable bills.</p>
<h2>Coding agents: the most important decision</h2>
<p>The tool you write code with defines your velocity more than any framework. The current landscape:</p>
<table>
<thead>
<tr>
<th>Tool</th>
<th>Base price</th>
<th>What stands out in 2026</th>
</tr>
</thead>
<tbody><tr>
<td>Claude Code</td>
<td>$20–200/mo</td>
<td>Dynamic Workflows: orchestrates dozens of parallel subagents from one session</td>
</tr>
<tr>
<td>Cursor</td>
<td>$20–200/mo</td>
<td>Composer 2.5 (in-house model) included in Pro; ~$4B ARR as of June</td>
</tr>
<tr>
<td>GitHub Copilot</td>
<td>$10–100/mo</td>
<td>Agent mode GA on VS Code and JetBrains; watch out for the new credit billing</td>
</tr>
<tr>
<td>Devin Desktop</td>
<td>$20/mo + usage</td>
<td>The former Windsurf, now Cognition&#39;s agent-management hub</td>
</tr>
<tr>
<td>Codex (OpenAI)</td>
<td>Bundled with ChatGPT</td>
<td>Open-source CLI + cloud tasks; 5M+ weekly users</td>
</tr>
<tr>
<td>Antigravity (Google)</td>
<td>Bundled with Gemini plans</td>
<td>Replaced Gemini CLI in June; async background workflows</td>
</tr>
</tbody></table>
<p>My read: the category no longer competes on &quot;better autocomplete&quot; — it competes on orchestration: how many agents you can direct at once, and with how much confidence. Claude Code is my daily driver for exactly that reason; Cursor is still the best editor if you want a traditional IDE with an agent inside.</p>
<blockquote>
<p>💡 Trap of the month: GitHub swapped Premium Requests for &quot;AI Credits&quot; in June, and teams are reporting bills 10x higher on agentic workflows. Whatever your tool, set spend alerts before you unleash parallel agents.</p>
</blockquote>
<h2>Models: July 2026 prices</h2>
<p>If you&#39;re building product on top of LLMs, this is what a million tokens costs today (input/output):</p>
<table>
<thead>
<tr>
<th>Model</th>
<th>Input</th>
<th>Output</th>
<th>Note</th>
</tr>
</thead>
<tbody><tr>
<td>Claude Fable 5</td>
<td>$10</td>
<td>$50</td>
<td>Anthropic&#39;s top of the line; 1M context</td>
</tr>
<tr>
<td>Claude Opus 4.8</td>
<td>$5</td>
<td>$25</td>
<td>The workhorse for code</td>
</tr>
<tr>
<td>Claude Sonnet 5</td>
<td>$2</td>
<td>$10</td>
<td>Intro pricing through August; then $3/$15</td>
</tr>
<tr>
<td>GPT-5.5</td>
<td>$5</td>
<td>$30</td>
<td>1M context</td>
</tr>
<tr>
<td>GPT-5.6 (Sol/Terra/Luna)</td>
<td>$1–5</td>
<td>$6–30</td>
<td>Released literally today</td>
</tr>
<tr>
<td>Gemini 3.1 Pro</td>
<td>$2</td>
<td>$12</td>
<td>Up to 200K context; more above that</td>
</tr>
<tr>
<td>DeepSeek V4 / Qwen 3.5</td>
<td>~open source</td>
<td>—</td>
<td>If you can self-host, the cost math changes leagues</td>
</tr>
</tbody></table>
<p>The sensible strategy for a SaaS: a cheap model (Sonnet 5, Terra, Gemini) for 90% of calls, escalating to the top tier only when the case justifies it. A 30-line model router saves you thousands of dollars a month.</p>
<h2>Web frameworks: less drama than it seems</h2>
<ul>
<li><strong>Next.js 16.2</strong> remains the rational default for SaaS: Turbopack is stable and on by default, React Compiler is integrated, and the ecosystem of examples is unbeatable. Boring and correct.</li>
<li><strong>Astro</strong> — just acquired by Cloudflare in January — is still my pick for anything content-shaped (this blog runs on Astro). Still MIT, still open governance.</li>
<li><strong>TanStack Start</strong> hit stable v1.0 in March and is the serious alternative if you want extreme type-safety without Next&#39;s magic.</li>
<li><strong>SvelteKit and React Router 7</strong> are mature and excellent; choosing them is more a matter of team taste than capability.</li>
</ul>
<p>The uncomfortable truth: with agents writing most of the code, framework choice matters less than it did three years ago. Agents perform better on frameworks with a bigger public corpus — another point for Next and Astro.</p>
<h2>Backend and infra: the stat of the year</h2>
<p>The stat that best sums up 2026: Supabase reported that the majority of its new databases are now deployed by AI agents, not humans — database creation growing 600% year over year. Your infrastructure is no longer chosen only by your team; agents &quot;choose&quot; it too, and they go with what they know how to use.</p>
<ul>
<li><strong>Postgres wins by a landslide</strong>: Supabase (just raised $500M at a $10.5B valuation) as a full backend, or Neon (now inside Databricks) if you only want the database.</li>
<li><strong>Vercel</strong> for deployment if you go with Next: Fluid Compute genuinely eliminated cold starts.</li>
<li><strong>Cloudflare Workers + D1 + R2</strong> is the cost-effective alternative — and with Astro in the family, increasingly polished.</li>
<li><strong>Runtimes</strong>: Node 24 is still the enterprise default. Bun — acquired by Anthropic in December — completed its Rust rewrite in May and is my pick for new projects: speed plus built-in tooling.</li>
</ul>
<h2>Your product&#39;s AI layer</h2>
<p>This is where I see the most over-engineering. What a startup actually needs:</p>
<pre><code class="language-text">// The minimum viable AI layer in 2026
// 1. SDK: Vercel AI SDK 6 (TypeScript), or the Claude Agent SDK
//    if the product IS an agent. LangGraph 1.0 if you need
//    complex state graphs (that&#39;s how Uber and Klarna use it).
// 2. Context: MCP. It&#39;s a Linux Foundation standard now,
//    97M monthly downloads. Don&#39;t invent your own protocol.
// 3. Vectors: pgvector up to ~10M vectors (~$30/mo).
//    Don&#39;t pay for a dedicated vector DB before you have the problem.
// 4. Evals from day one: if you don&#39;t measure LLM quality,
//    every deploy is a gamble.
</code></pre>
<p>Point 3 deserves emphasis: pgvector in your regular Postgres covers almost any startup with 8–25ms latencies. Notion cut its search costs ~60% by moving off a dedicated vector DB onto object-storage-backed search; you probably don&#39;t even need that yet.</p>
<h2>Vibe coding: use it for what it is</h2>
<p>Lovable ($400M ARR in February, reportedly in talks to raise at an eleven-figure valuation), Replit ($9B valuation in March), v0, Bolt. They are idea-validation machines: prompt to deployed prototype in an afternoon. What they are not — yet — is the foundation of a product that has to scale with serious security and data requirements. The pattern I see working: prototype in Lovable or v0, validate with real users, and rebuild on your stack once there&#39;s signal.</p>
<h2>My concrete recipe</h2>
<p>If I started a B2B SaaS tomorrow, with no more context than &quot;I want to validate fast and not burn money&quot;:</p>
<ul>
<li><strong>Code:</strong> Claude Code as the primary agent + human review of everything touching money, permissions or data.</li>
<li><strong>App:</strong> Next.js 16 + TypeScript on Vercel. <strong>Content/marketing:</strong> Astro.</li>
<li><strong>Data:</strong> Supabase (Postgres + auth + storage + pgvector).</li>
<li><strong>AI:</strong> Vercel AI SDK 6, Sonnet 5 by default with escalation to Opus 4.8, MCP for integrations, evals with real cases from week one.</li>
<li><strong>Runtime:</strong> Bun locally and in CI; Node 24 wherever the platform demands it.</li>
</ul>
<p>Is it the most powerful stack possible? No. It&#39;s the one that lets you iterate every single day with two and a half people — and in 2026 that&#39;s exactly the game: the edge is no longer in the infrastructure you set up, but in how fast you learn what to build.</p>
]]></content:encoded>
    </item>
    <item>
      <title>AI-built startups in 2026: the real numbers behind the hype</title>
      <link>https://yohangel.com/en/blog/startups-ia-casos-exito-2026/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/startups-ia-casos-exito-2026/</guid>
      <description>Cursor, Lovable, Cognition, Base44, Cal AI and the YC data: what they actually earn, with how many people, and which patterns repeat. With charts.</description>
      <pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>IA</category>
      <content:encoded><![CDATA[<p><strong>Everyone talks about AI hype every day; almost nobody talks about the numbers. So I gathered the public figures — verified against financial press, official blogs and confirmed funding rounds — for 2026 so far. The short conclusion: the growth curves we&#39;re seeing have never existed before in the history of software.</strong></p>
<p>Before we start, a note on method: I only use publicly sourced figures. When a number is a third-party estimate (rather than company-confirmed) I flag it as such. Even with that filter, what remains is hard to believe.</p>
<h2>The curve that defines the era: Cursor</h2>
<p>Anysphere (Cursor) was at one point the fastest SaaS in history to reach $100M ARR. That was January 2025. What happened next is the chart that defines this era:</p>
<svg viewBox="0 0 700 300" width="100%" role="img" aria-label="Cursor ARR: from 100 million in January 2025 to 4 billion in June 2026" style="margin:0 0 22px; font-family: var(--font-mono);">
  <text x="20" y="24" fill="var(--color-ink)" font-size="14" font-weight="700">Cursor — ARR in $ millions</text>
  <line x1="30" y1="250" x2="670" y2="250" stroke="var(--hairline)" stroke-width="1" />
  <rect x="45" y="245" width="66" height="5" rx="2" fill="var(--color-violet)" />
  <rect x="150" y="225" width="66" height="25" rx="3" fill="var(--color-violet)" />
  <rect x="255" y="200" width="66" height="50" rx="4" fill="var(--color-violet)" />
  <rect x="360" y="150" width="66" height="100" rx="4" fill="var(--color-violet)" />
  <rect x="465" y="100" width="66" height="150" rx="4" fill="var(--color-violet-light)" />
  <rect x="570" y="50" width="66" height="200" rx="4" fill="var(--color-violet-light)" />
  <text x="78" y="238" fill="var(--color-body)" font-size="12" text-anchor="middle">100</text>
  <text x="183" y="218" fill="var(--color-body)" font-size="12" text-anchor="middle">500</text>
  <text x="288" y="193" fill="var(--color-body)" font-size="12" text-anchor="middle">1,000</text>
  <text x="393" y="143" fill="var(--color-body)" font-size="12" text-anchor="middle">2,000</text>
  <text x="498" y="93" fill="var(--color-body)" font-size="12" text-anchor="middle">3,000</text>
  <text x="603" y="43" fill="var(--color-ink)" font-size="13" font-weight="700" text-anchor="middle">4,000</text>
  <text x="78" y="270" fill="var(--color-faint)" font-size="11" text-anchor="middle">Jan 25</text>
  <text x="183" y="270" fill="var(--color-faint)" font-size="11" text-anchor="middle">Jun 25</text>
  <text x="288" y="270" fill="var(--color-faint)" font-size="11" text-anchor="middle">Nov 25</text>
  <text x="393" y="270" fill="var(--color-faint)" font-size="11" text-anchor="middle">Feb 26</text>
  <text x="498" y="270" fill="var(--color-faint)" font-size="11" text-anchor="middle">Apr 26</text>
  <text x="603" y="270" fill="var(--color-faint)" font-size="11" text-anchor="middle">Jun 26</text>
</svg><p>From $100M to ~$4B in annualized revenue in 17 months, with around $2.6B coming from enterprise (B2B). Its November 2025 Series D valued it at $29.3B, and by April 2026 there were reports of talks to raise at ~$50B. For calibration: Salesforce took more than a decade to reach the revenue Cursor hit in three years.</p>
<h2>The multipliers of 2026</h2>
<p>Cursor is not an isolated case. This is the table for the year so far:</p>
<table>
<thead>
<tr>
<th>Company</th>
<th>~1 year ago</th>
<th>Now (2026)</th>
<th>Valuation</th>
</tr>
</thead>
<tbody><tr>
<td>Cognition (Devin)</td>
<td>$37M ARR (May 25)</td>
<td>$492M ARR (May 26) — 13x</td>
<td>$26B ($1B round, May 26)</td>
</tr>
<tr>
<td>Lovable</td>
<td>$100M ARR (~Jul 25)</td>
<td>$400M ARR (Feb 26); ~$500M est.</td>
<td>$6.6B (Dec 25); reported talks at ~$12B</td>
</tr>
<tr>
<td>ElevenLabs</td>
<td>$3.3B val. (Jan 25)</td>
<td>~$500M ARR est.</td>
<td>$11B (Series D, Feb 26)</td>
</tr>
<tr>
<td>Harvey (legal)</td>
<td>$100M ARR (Aug 25)</td>
<td>$300M+ ARR (Jun 26)</td>
<td>$11B (Mar 26)</td>
</tr>
<tr>
<td>Mercor</td>
<td>—</td>
<td>$1B ARR (early 26), profitable</td>
<td>$10B (Oct 25)</td>
</tr>
<tr>
<td>Replit</td>
<td>$150M annualized (Sep 25)</td>
<td>~$525M est. (Apr 26)</td>
<td>$9B (Mar 26)</td>
</tr>
</tbody></table>
<p>And the backdrop: Anthropic went from a $9B run-rate at the end of 2025 to $47B in May 2026, and Claude Code alone hit $1B annualized within six months of launch. Demand for building with AI is the tide lifting all of these boats.</p>
<h2>The most disruptive stat: revenue per employee</h2>
<p>What truly breaks mental models is not how much they earn, but with how few people. Annual revenue per employee, in millions of dollars:</p>
<svg viewBox="0 0 700 210" width="100%" role="img" aria-label="Revenue per employee: Cursor 13 million, Midjourney 4.7, Lovable 2, Gamma 2" style="margin:0 0 22px; font-family: var(--font-mono);">
  <text x="20" y="24" fill="var(--color-ink)" font-size="14" font-weight="700">Annual revenue per employee ($M)</text>
  <text x="160" y="62" fill="var(--color-body)" font-size="12" text-anchor="end">Cursor</text>
  <rect x="172" y="48" width="440" height="20" rx="4" fill="var(--color-violet-light)" />
  <text x="622" y="62" fill="var(--color-ink)" font-size="12" font-weight="700">~13</text>
  <text x="160" y="100" fill="var(--color-body)" font-size="12" text-anchor="end">Midjourney</text>
  <rect x="172" y="86" width="159" height="20" rx="4" fill="var(--color-violet)" />
  <text x="341" y="100" fill="var(--color-body)" font-size="12">4.7</text>
  <text x="160" y="138" fill="var(--color-body)" font-size="12" text-anchor="end">Lovable</text>
  <rect x="172" y="124" width="68" height="20" rx="4" fill="var(--color-sky)" />
  <text x="250" y="138" fill="var(--color-body)" font-size="12">~2</text>
  <text x="160" y="176" fill="var(--color-body)" font-size="12" text-anchor="end">Gamma</text>
  <rect x="172" y="162" width="68" height="20" rx="4" fill="var(--color-mint)" />
  <text x="250" y="176" fill="var(--color-body)" font-size="12">~2</text>
</svg><p>For perspective: a well-run traditional software company sits around $200,000–400,000 per employee. The cases above are 5 to 40 times higher:</p>
<ul>
<li><strong>Gamma</strong> passed $100M ARR with about 50 people, profitable for over two years.</li>
<li><strong>Midjourney</strong> earns ~$500M with fewer than 110 employees and has never raised a dollar of venture capital. Ever.</li>
<li><strong>Mercor</strong> crossed $1B ARR with ~200 people and positive cash flow.</li>
<li><strong>Cursor</strong> operates ~$4B with around 300 employees (an estimate; the figure isn&#39;t official).</li>
</ul>
<h2>The small ones: solo founders who exited through the front door</h2>
<p>Unicorns get the headlines, but the more replicable stories are elsewhere:</p>
<ul>
<li><strong>Base44 (Maor Shlomo).</strong> A single bootstrapped founder, building with AI. Launched in early 2025, hit $1.5M ARR in 4 weeks and $3.5M with 300,000 users in 6 months. Wix bought it for <strong>$80M in cash</strong> six months after launch — and subsequent milestones paid him an extra $38M. Eight employees at the time of sale.</li>
<li><strong>Cal AI (Zach Yadegari and Henry Langmack).</strong> Two founders who started at 17: an app that counts calories from photos using AI. $30M in revenue in 2025, $5.7M in January 2026 alone, and in March — at $50M ARR — MyFitnessPal acquired it.</li>
</ul>
<blockquote>
<p>💡 The pattern they share: neither invented new technology. They applied models anyone can call over an API to a boring, concrete problem, and executed faster than everyone else. The edge is no longer access to AI — it&#39;s iteration speed on a use case.</p>
</blockquote>
<h2>And underneath it all: AI is already writing the code</h2>
<p>These successes float on a measurable structural shift. Google has been publishing it with dates:</p>
<svg viewBox="0 0 700 260" width="100%" role="img" aria-label="Share of Google's new code generated by AI: 25 percent in 2024, 50 by late 2025, 75 in April 2026" style="margin:0 0 22px; font-family: var(--font-mono);">
  <text x="20" y="24" fill="var(--color-ink)" font-size="14" font-weight="700">Google — % of new code generated by AI</text>
  <line x1="30" y1="215" x2="670" y2="215" stroke="var(--hairline)" stroke-width="1" />
  <rect x="90" y="160" width="120" height="55" rx="4" fill="var(--color-sky)" />
  <rect x="290" y="105" width="120" height="110" rx="4" fill="var(--color-violet)" />
  <rect x="490" y="50" width="120" height="165" rx="4" fill="var(--color-violet-light)" />
  <text x="150" y="150" fill="var(--color-body)" font-size="13" text-anchor="middle">25%</text>
  <text x="350" y="95" fill="var(--color-body)" font-size="13" text-anchor="middle">50%</text>
  <text x="550" y="40" fill="var(--color-ink)" font-size="14" font-weight="700" text-anchor="middle">75%</text>
  <text x="150" y="237" fill="var(--color-faint)" font-size="11" text-anchor="middle">early 2024</text>
  <text x="350" y="237" fill="var(--color-faint)" font-size="11" text-anchor="middle">late 2025</text>
  <text x="550" y="237" fill="var(--color-faint)" font-size="11" text-anchor="middle">April 2026</text>
</svg><p>At Anthropic the reported internal figure exceeds 90%. And at Y Combinator — the best window into the immediate future: already in the W25 batch, 25% of startups had 95% of their code AI-generated; at the W26 Demo Day (March 2026), <strong>14 startups arrived with over $1M ARR before Demo Day</strong> — triple the previous year and an all-time record — with average growth of 14% per week, the highest YC has recorded in 20 years.</p>
<h2>The patterns that repeat</h2>
<p>After looking at all these cases, the common denominators I see:</p>
<ul>
<li><strong>Tiny teams by design, not for lack of money.</strong> Lovable raised hundreds of millions and stays at ~250 people. The headcount constraint is strategy: less coordination, more agents.</li>
<li><strong>Speed as the only moat.</strong> Almost none of them have defensible technology per se; they have a shipping cadence competitors can&#39;t match.</li>
<li><strong>Distribution first.</strong> Cal AI grew through social media before the product matured; Lovable turned every user into marketing with &quot;look what I built.&quot;</li>
<li><strong>B2B pays for the curves.</strong> Cursor&#39;s and Devin&#39;s acceleration comes from enterprise contracts (Goldman Sachs, Citi and NASA are Cognition customers), not individual developers.</li>
</ul>
<h2>The fine print</h2>
<p>To be fair to the data: several 2026 ARR figures (Lovable ~$500M, Replit ~$525M, Claude Code) are estimates from trackers like Sacra, not official confirmations; Lovable&#39;s round at ~$12B was still in negotiation as of this writing; and for every success story there are hundreds of GPT wrappers that died without a headline — survivorship bias in this list is maximal.</p>
<p>But even discounting all of that, the signal is clear: in 2026 there is no longer a correlation between team size and business size. And for anyone who knows how to build software, that is the best news in decades: it has never been this cheap to try.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Fine-tuning, RAG or prompting: how I decide which to use</title>
      <link>https://yohangel.com/en/blog/fine-tuning-vs-rag-decision/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/fine-tuning-vs-rag-decision/</guid>
      <description>An honest decision tree for choosing between fine-tuning a model, building RAG, or staying with prompting, with the real cost of each branch.</description>
      <pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>IA</category>
      <content:encoded><![CDATA[<p><strong>Almost every time someone tells me &quot;we need fine-tuning,&quot; what they actually need is better context. Fine-tuning is the first word out because it sounds like the serious solution, the one that &quot;teaches&quot; the model. But in product I&#39;ve changed very few behaviors with fine-tuning and a huge number with a good prompt or with well-built RAG. The rule that has worked for me is to start cheap and reversible, and only climb a rung when the previous one proves to be capped. This article is that decision tree, with the trade-offs I accept on each branch and the signals that tell me when I&#39;ve chosen wrong.</strong></p>
<h2>First I ask: is this a knowledge problem or a behavior problem?</h2>
<p>This is the fork that orders everything else. If the model fails because it doesn&#39;t know something —your domain data, internal documentation, facts after its training cutoff— it&#39;s a knowledge problem, and knowledge is injected through context, not baked into weights. There, RAG is the default answer.</p>
<p>If the model knows what it needs but consistently responds with the wrong format, tone, or structure, it&#39;s a behavior problem. And behavior is a candidate for fine-tuning… but only after prompting falls short, which is later than people think.</p>
<p>Confusing these two branches is the most expensive mistake I see. Nobody fixes the model not knowing your product catalog with fine-tuning: you train it today and tomorrow the catalog changed. And nobody fixes the model refusing to return clean JSON with RAG: no matter how many documents you feed it, the problem is shape, not data.</p>
<h2>Prompting: the rung I almost always underestimate</h2>
<p>My starting point is always prompting, because it&#39;s the only change I ship in minutes and revert in seconds. Before touching anything else, I squeeze clear instructions, representative few-shot examples, and an explicit output structure. At JXBS, a good chunk of what looked like it needed a fine-tuned model was solved with three well-chosen examples in the prompt and a strict output schema.</p>
<p>The cost of this branch is context: every example you add takes tokens you pay for on every call, and a 20-example prompt becomes expensive and slow at scale. That&#39;s exactly the symptom that maybe it&#39;s time to climb a rung: when you need so many examples that the prompt is half your bill, the pattern is already there and it may be worth baking in.</p>
<blockquote>
<p>💡 If you haven&#39;t tried to solve it with prompting for at least an afternoon, you don&#39;t have the data to justify fine-tuning. &quot;We tried a prompt and it didn&#39;t work&quot; isn&#39;t prompting, it&#39;s an anecdote.</p>
</blockquote>
<h2>RAG: for knowledge that changes and must be citable</h2>
<p>When the problem is knowledge, RAG almost always wins for a reason beyond the technical one: traceability. A RAG system can tell you <em>which document</em> the answer came from. A fine-tuned model gives you the answer from inside its weights, with no source, and when it hallucinates you have nowhere to look.</p>
<pre><code class="language-typescript">// Knowledge lives outside the model and updates without retraining
const context = await searchChunks(question, { topK: 5 });
const answer = await llm.complete({
  system: &#39;Answer only with information in &lt;context&gt;. If it is not there, say so.&#39;,
  context,
  question,
});
</code></pre>
<p>The trade-off I accept with RAG is operational: now I maintain an ingestion pipeline, embeddings, a vector index, and a chunking strategy. It&#39;s more infrastructure than a prompt. In exchange, I update knowledge by changing documents, not retraining, and in a fast-moving domain that&#39;s worth gold.</p>
<h2>Fine-tuning: the last rung, and I know why I&#39;m climbing it</h2>
<p>I reach fine-tuning only when three things hold at once: the behavior I want is stable and repeated, prompting gives it to me but at an unsustainable token cost, and I have enough quality real examples to train on. If any of the three is missing, I don&#39;t climb.</p>
<p>What I buy with fine-tuning is a model that does what I want with a short prompt, cheaper and faster per call. What I pay is rigidity: every behavior change is a retrain, the dataset becomes an asset you have to version and care for, and I&#39;ve introduced an artifact that can go stale silently. Fine-tuning doesn&#39;t eliminate RAG either: the most powerful thing I&#39;ve built combines a model fine-tuned on <em>how</em> to answer with RAG for <em>what</em> to know.</p>
<h2>The mistake of skipping rungs</h2>
<p>The antipattern that&#39;s cost me most is starting at the end. Teams that build fine-tuning for a problem a prompt would have solved, and end up with an expensive-to-maintain model that performs no better than the baseline they never measured. Starting cheap isn&#39;t being conservative: it&#39;s generating the evidence that justifies the next step. When you finally do fine-tuning after exhausting prompting and RAG, you know exactly what you&#39;re buying and why. That clarity is the difference between an engineering decision and a purchase driven by hype.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Step Functions: when an orchestrator beats chaining Lambdas</title>
      <link>https://yohangel.com/en/blog/step-functions-orquestacion/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/step-functions-orquestacion/</guid>
      <description>Why I move the coordination logic of long processes out of code and into an explicit state machine, and the trade-offs that implies.</description>
      <pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>AWS</category>
      <content:encoded><![CDATA[<p><strong>For a while my way of chaining steps in AWS was the obvious one: a Lambda does its part and on finishing invokes the next, which invokes the next. It works with two steps. With six, with retries, with a step that sometimes takes ten minutes, and with the need to know where a process that failed last night got stuck, that approach becomes a tangle impossible to debug. Step Functions changed that for me: I pulled the coordination logic out of the code and into an explicit state machine. I chose this model in several flows because the cost of not knowing where each execution was outweighed the alternatives; in exchange I accepted learning a new definition language and paying per state transition. This article is about when it pays off and when it doesn&#39;t.</strong></p>
<h2>The problem: orchestration hidden in the code</h2>
<p>When one Lambda invokes the next, the process logic —the order, the retries, what to do if step three fails— lives scattered across functions. There&#39;s no single place to read &quot;this is how the whole flow works.&quot; To understand it you open five files and reconstruct the graph in your head.</p>
<p>The day that process fails in production halfway through, the key question is &quot;which step did it get stuck on and with what data?&quot;. With chained Lambdas, the answer is buried in logs from five different functions that you have to correlate by hand. With an orchestrator, the answer is a screen: you see the execution, the exact state where it died, and the input it received.</p>
<h2>The state machine as executable documentation</h2>
<p>What I value most about Step Functions is that the flow definition <em>is</em> the diagram. There&#39;s no Confluence doc going stale: the source of truth for the process is the state machine running in production.</p>
<pre><code class="language-json">{
  &quot;StartAt&quot;: &quot;ValidateOrder&quot;,
  &quot;States&quot;: {
    &quot;ValidateOrder&quot;: {
      &quot;Type&quot;: &quot;Task&quot;,
      &quot;Resource&quot;: &quot;arn:aws:lambda:...:validate&quot;,
      &quot;Retry&quot;: [{ &quot;ErrorEquals&quot;: [&quot;Timeout&quot;], &quot;MaxAttempts&quot;: 3 }],
      &quot;Next&quot;: &quot;ChargePayment&quot;
    },
    &quot;ChargePayment&quot;: {
      &quot;Type&quot;: &quot;Task&quot;,
      &quot;Resource&quot;: &quot;arn:aws:lambda:...:charge&quot;,
      &quot;Catch&quot;: [{ &quot;ErrorEquals&quot;: [&quot;States.ALL&quot;], &quot;Next&quot;: &quot;Compensate&quot; }],
      &quot;Next&quot;: &quot;Confirm&quot;
    },
    &quot;Compensate&quot;: { &quot;Type&quot;: &quot;Task&quot;, &quot;Resource&quot;: &quot;arn:aws:lambda:...:revert&quot;, &quot;End&quot;: true },
    &quot;Confirm&quot;: { &quot;Type&quot;: &quot;Task&quot;, &quot;Resource&quot;: &quot;arn:aws:lambda:...:confirm&quot;, &quot;End&quot;: true }
  }
}
</code></pre>
<p>Notice where the retry and compensation policies live: in the definition, declarative, not scattered across <code>try/catch</code> in every function. Each Lambda goes back to doing a single thing and the orchestrator handles the rest. That&#39;s the split of responsibilities I was after.</p>
<h2>Retries and compensation without hand-writing them</h2>
<p>The reason I migrate flows to Step Functions most often is failure handling. In chained Lambdas, retrying with backoff, distinguishing transient from permanent errors, and compensating already-executed steps when something fails halfway is code you write, test, and maintain yourself, badly, in every function.</p>
<p>In the state machine, <code>Retry</code> and <code>Catch</code> are part of the definition. A timeout retries three times with exponential backoff without me writing a loop. And the saga pattern —if payment fails after reserving inventory, revert the reservation— is expressed with a compensation state you jump to with <code>Catch</code>, instead of with flags scattered across half the backend.</p>
<blockquote>
<p>💡 The question I use to decide: do I need to know the exact step where a failed execution got stuck, weeks later? If the answer is yes, explicit orchestration pays for itself in the first incident.</p>
</blockquote>
<h2>Where I do NOT use Step Functions</h2>
<p>Being honest about the trade-offs is what separates a decision from an act of faith. I don&#39;t put Step Functions in everything.</p>
<p>For a simple synchronous flow where the user waits for a response on screen, an orchestrator only adds latency and a layer nobody asked for; a direct Lambda is right. I also don&#39;t use it for very high-volume, short-duration processes where every state transition is billed: the Standard model charges per transition and at millions of short executions that scales into your bill in surprising ways. For that case I look at Express, which changes the pricing model and the durability guarantee, or straight to an SQS queue with a consumer that does all the work at once.</p>
<p>And there&#39;s a human cost: the definition language is one more thing the team has to learn and read. If your flow has two steps and will never grow, that curve isn&#39;t worth it.</p>
<h2>The real price I paid</h2>
<p>The most expensive part of adopting Step Functions wasn&#39;t the service, it was rewiring the team&#39;s mindset. We went from &quot;the logic is in the code&quot; to &quot;the logic is in the state definition,&quot; and for a while people kept putting coordination inside the Lambdas out of habit, duplicating what the machine already did. The rule I set was clear: Lambdas do work, the state machine decides the order. Once the team internalized that boundary, debugging a long process stopped being log archaeology and became looking at a diagram that tells you exactly where you are. That shift, not the JSON syntax, is what makes it worth it.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Web Workers: getting heavy work off the main thread</title>
      <link>https://yohangel.com/en/blog/web-workers-hilo-principal/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/web-workers-hilo-principal/</guid>
      <description>When moving computation to a real worker improves the UI and when it only adds complexity, with the pattern I use to avoid fighting postMessage.</description>
      <pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>Frontend</category>
      <content:encoded><![CDATA[<p><strong>The browser&#39;s main thread does too many things: it runs your JavaScript, computes layout, paints, and responds to clicks. All in the same queue. When you drop a heavy computation in there —parsing a CSV with thousands of rows, filtering a large dataset, processing an image— you block that thread and the interface freezes: scrolling stutters, buttons don&#39;t respond, the spinner doesn&#39;t even spin because the thread that would animate it is busy. Web Workers exist for that: to move that computation to another thread and give the main one back its single important responsibility, which is keeping the UI alive. I&#39;ve used workers in several PWAs and the improvement in perception is real, but I&#39;ve also seen them dropped in where they didn&#39;t belong. This article is about when they&#39;re worth it.</strong></p>
<h2>The symptom that tells you you need a worker</h2>
<p>Not every computation goes to a worker. The concrete signal I look for is a <em>long task</em>: a synchronous function that occupies the main thread for more than a few dozen milliseconds and during which the page stops responding. If you open the performance panel and see long yellow blocks right when the UI stutters, there&#39;s your culprit.</p>
<p>Typical cases that have led me to a worker: parsing and transforming large files the user uploads, sorting or filtering a dataset of tens of thousands of rows client-side, geometry or image calculations, and running light models in the browser. What they have in common is that they&#39;re pure, prolonged CPU work, not network waits. For waiting on a server you don&#39;t need a worker: <code>async/await</code> already frees the thread. The worker is for work that <em>burns CPU</em>, not work that <em>waits</em>.</p>
<h2>The cost people forget: the message boundary</h2>
<p>A worker doesn&#39;t share memory with the main thread. They communicate by passing messages, and that data is <em>copied</em> when crossing the boundary (structured clone). That detail is what decides whether the worker pays off: if you send a huge object to the worker, wait for a trivial computation, and get another huge object back, the cost of serializing and copying eats the gain.</p>
<p>The rule I follow is to push <em>a lot</em> of work per boundary crossing. A worker pays off when you send it data once and it does a big computation, not when you call it in a tight loop for small operations. When the data is genuinely large, I use transferable objects: an <code>ArrayBuffer</code> is <em>transferred</em> instead of copied, changing owner without duplicating memory.</p>
<pre><code class="language-typescript">// Main thread: transfer the buffer, don&#39;t copy it
const worker = new Worker(new URL(&#39;./process.ts&#39;, import.meta.url), { type: &#39;module&#39; });
worker.postMessage({ buffer }, [buffer]); // the second arg transfers ownership
worker.onmessage = (e) =&gt; renderResult(e.data);
</code></pre>
<h2>The pattern I use to avoid suffering with postMessage</h2>
<p>The raw <code>postMessage</code> / <code>onmessage</code> API turns into event spaghetti the moment you have more than two message types. What I do is wrap the worker in a promise, so that from the rest of the code calling the worker feels like a normal <code>await</code> and not like wiring an event bus.</p>
<pre><code class="language-typescript">function runInWorker&lt;T&gt;(worker: Worker, payload: unknown): Promise&lt;T&gt; {
  return new Promise((resolve, reject) =&gt; {
    worker.onmessage = (e) =&gt; resolve(e.data as T);
    worker.onerror = (e) =&gt; reject(e);
    worker.postMessage(payload);
  });
}
</code></pre>
<p>For serious cases I use Comlink, which does exactly this but properly: it lets you call worker functions as if they were local, with <code>await</code>, and hides all the messaging. The first time you wrap a worker like this, the mental barrier of &quot;this is complicated&quot; disappears: the worker becomes just another async function.</p>
<blockquote>
<p>💡 A worker doesn&#39;t make your code faster; it makes your UI <em>smoother</em>. The computation takes the same time or slightly more because of the data copy. What you gain is that the user can keep scrolling and clicking meanwhile. It&#39;s an improvement in perception, not throughput.</p>
</blockquote>
<h2>Where I do NOT reach for a worker</h2>
<p>Being honest about the trade-off matters. I don&#39;t use a worker when the computation is short: if it takes 5 ms, moving it to another thread only adds messaging latency and a layer of complexity for nothing. Nor when the work is waiting on the network, because that no longer blocks. And I avoid the worker if the algorithm needs to touch the DOM: workers have no DOM access by design, and if your &quot;computation&quot; actually manipulates the interface, it&#39;s not a candidate.</p>
<p>The other cost is tooling and maintenance: a worker is another file, another entry point for the bundler, another piece someone has to understand when reading the code. With Vite or whatever modern bundler, the friction dropped a lot, but it&#39;s not zero. That&#39;s why my threshold is clear: a worker only when there&#39;s a measurable long task that visibly stalls the UI. If I can&#39;t point to that yellow block in the profiler, there&#39;s no worker. Optimizing a block nobody notices is adding complexity to show off architecture, and that isn&#39;t engineering, it&#39;s decoration.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Semantic caching for LLMs: pay once for answers you already gave</title>
      <link>https://yohangel.com/en/blog/cache-semantico-llm/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/cache-semantico-llm/</guid>
      <description>How to use embeddings to detect repeated questions and serve cached answers without calling the model, pitfalls included.</description>
      <pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>IA</category>
      <content:encoded><![CDATA[<p><strong>Your LLM bill in production doesn&#39;t grow because of hard questions: it grows because of repeated ones. In any product with an assistant, a huge share of the traffic is variations of the same queries — &quot;how do I change my password?&quot;, &quot;how to reset my password&quot;, &quot;forgot my password&quot;. A traditional exact-key cache captures none of that, because the text is never identical. A semantic cache does: it turns the question into an embedding, checks whether you already answered something similar enough, and if the nearest neighbor clears a similarity threshold, it returns the stored answer without touching the model. I chose this pattern in a real product because cost per request mattered more than absolute freshness of every answer; in exchange, I accepted the complexity of managing thresholds and invalidating entries. This article is what I learned.</strong></p>
<h2>The mechanics: pgvector is enough</h2>
<p>You don&#39;t need new infrastructure. If you already use PostgreSQL with pgvector for semantic search — which is my usual setup with Prisma — the cache is just one more table:</p>
<pre><code class="language-sql">CREATE TABLE llm_cache (
  id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
  query_text TEXT NOT NULL,
  query_embedding vector(1536) NOT NULL,
  response TEXT NOT NULL,
  model TEXT NOT NULL,
  hit_count INT DEFAULT 0,
  created_at TIMESTAMPTZ DEFAULT now(),
  expires_at TIMESTAMPTZ NOT NULL
);

CREATE INDEX ON llm_cache
  USING hnsw (query_embedding vector_cosine_ops);
</code></pre>
<p>The read path is a nearest-neighbor query with a threshold:</p>
<pre><code class="language-typescript">const [hit] = await prisma.$queryRaw&lt;CacheHit[]&gt;`
  SELECT response, 1 - (query_embedding &lt;=&gt; ${embedding}::vector) AS similarity
  FROM llm_cache
  WHERE expires_at &gt; now()
  ORDER BY query_embedding &lt;=&gt; ${embedding}::vector
  LIMIT 1
`;

if (hit &amp;&amp; hit.similarity &gt;= 0.92) {
  return hit.response; // cache hit: zero tokens spent
}
</code></pre>
<p>If there&#39;s no hit, you call the model, store the response with its embedding, and the next user asking the same thing in different words no longer pays.</p>
<h2>The threshold is a product decision, not a technical detail</h2>
<p>The whole pattern lives or dies on that <code>0.92</code>. Too low, and you serve wrong answers to questions that only look alike on the surface — &quot;how do I cancel my subscription?&quot; and &quot;how do I cancel a payment?&quot; have uncomfortably close embeddings. Too high, and your hit rate collapses and the cache never pays for its own complexity.</p>
<p>My approach: start at 0.95, log every hit with both the original and the cached question, and manually review a sample every week. Lower the threshold only when the false positives in the next band are acceptable. In domains with a closed vocabulary (support for a specific product) I&#39;ve gone down to 0.90; in open domains I wouldn&#39;t go below 0.93.</p>
<blockquote>
<p>💡 The similarity threshold is not a hyperparameter you tune once: it&#39;s a product policy defining how much imprecision you tolerate in exchange for cost. Treat it as such — with logging, periodic review, and an owner.</p>
</blockquote>
<h2>Invalidation: the hidden price</h2>
<p>A semantic cache inherits the classic cache problem and adds one of its own. The classic one: answers expire. If your product changes its billing flow, every cached answer about billing is now a well-written lie. That&#39;s why every entry carries <code>expires_at</code> and, more importantly, a category-based invalidation mechanism: I tag entries with the functional area they touch, and when I deploy a change in that area, I wipe the whole category. It&#39;s blunt, but it&#39;s predictable.</p>
<p>The problem of its own: personalized answers. If the model&#39;s response includes user data (&quot;your current plan is Pro&quot;), caching it and serving it to another user is a data leak, not an optimization. My rule is binary: only what passes a prior &quot;generic question&quot; classifier enters the cache. Anything that depends on user context stays out, no exceptions and no special cases.</p>
<h2>When not to do it</h2>
<p>Trade-offs rule. A semantic cache pays off when traffic has high semantic redundancy (support, FAQs, onboarding), cost per call is relevant, and latency matters — a hit responds in ~50ms versus the 2-4 seconds of a large model. It doesn&#39;t pay off when every question is unique (creative tools, analysis of user documents), when answers depend on personal context, or when volume is so low that the LLM bill is noise.</p>
<p>In my case, the signal to build it came from looking at the logs: when you see the same intent expressed twenty different ways day after day, the cache justifies itself. If you haven&#39;t looked at your logs yet, start there — you may not need any of this, and that&#39;s good news too.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Container queries: the day I stopped lying to my components</title>
      <link>https://yohangel.com/en/blog/container-queries-componentes/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/container-queries-componentes/</guid>
      <description>Why media queries break the promise of a reusable component and how container queries actually keep it.</description>
      <pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>Frontend</category>
      <content:encoded><![CDATA[<p><strong>A reusable component promises something simple: drop it anywhere and it works. For years that promise was a lie, and the lie had a name: media query. A card that adapts with <code>@media (min-width: 768px)</code> doesn&#39;t react to its own space, it reacts to the window width. Put it in a narrow sidebar with the screen at 1400px and the card thinks it has room to spare, because it&#39;s looking at the wrong variable. Container queries fix this at the root: the component responds to the size of its container, not the viewport. I adopted this in a real product&#39;s design system and it saved me the most frustrating class of bug to maintain — the &quot;looks fine in the demo, broken in its real spot&quot; kind. This article is why the shift matters more than it seems.</strong></p>
<h2>The media query lie</h2>
<p>The problem is one of abstraction level. A well-built component knows nothing about the outside world: it takes props, renders, and is supposed to adapt to wherever you place it. But a media query breaks that encapsulation entirely, because it queries a global state —the viewport width— that the component neither controls nor should know about.</p>
<p>The result is that the same component needs to know which layout it&#39;s going to live in. The card that looks perfect in the main three-column grid breaks in the sidebar, not because it&#39;s poorly built, but because its breakpoint assumes a window width that doesn&#39;t match the width it actually has. You end up adding props like <code>variant=&quot;compact&quot;</code> to tell it by hand what it should be able to figure out on its own. Every one of those props is a leak in the abstraction.</p>
<h2>How it works: the container declares, the child queries</h2>
<p>The mechanism has two parts. First you declare that an element is a query container. Then, inside it, you query its size instead of the screen&#39;s.</p>
<pre><code class="language-css">.card-wrapper {
  container-type: inline-size;
  container-name: card;
}

.card {
  display: grid;
  grid-template-columns: 1fr;
}

@container card (min-width: 400px) {
  .card {
    grid-template-columns: 120px 1fr;
  }
}
</code></pre>
<p>What changes from a media query is subtle and huge at once: <code>@container card (min-width: 400px)</code> doesn&#39;t ask &quot;is the screen wide?&quot; but &quot;is my container at least 400px?&quot;. The same card, without a single extra prop, renders in one column inside a 300px sidebar and in two columns inside a wide grid. The component finally adapts to its reality, not to an assumption about the window.</p>
<h2>The case that convinced me: one card in three different places</h2>
<p>In the design system I had a <code>ProductCard</code> that appeared in three contexts: the catalog grid (wide), a recommendations carousel (medium), and a &quot;recently viewed&quot; sidebar (narrow). With media queries I had three variants and an <code>if</code> in the consumer deciding which to use based on where it went. Three code paths for a single card.</p>
<p>With container queries I deleted the variants. The <code>ProductCard</code> became a single implementation that queries its container and reorganizes on its own. The component stopped needing context about its location, which is exactly what a component should not need. Fewer props, fewer branches, less surface where something can drift out of sync.</p>
<blockquote>
<p>💡 If your component needs a <code>variant</code> prop just to fit different widths, you don&#39;t have an API design problem: you have a media query where a container query should be. Fix it at the bottom and the prop disappears by itself.</p>
</blockquote>
<h2>Trade-offs I accepted</h2>
<p>None of this is free. Declaring <code>container-type: inline-size</code> creates a size containment context, and that has a consequence you have to understand: the container stops sizing itself to its content along the queried axis. In practice this rarely bites if you wrap with a dedicated wrapper instead of turning an element that already did other things into a container, but it&#39;s a real footnote, not a detail you can ignore.</p>
<p>I also accepted a bit more nesting in the markup. The clean pattern is a <code>wrapper</code> that is the container and a child that is what gets styled; querying an element&#39;s size while also styling that same element with the query doesn&#39;t work well. One extra div per component is a price I pay gladly in exchange for deleting a whole category of context props.</p>
<h2>Why this is more than a convenience</h2>
<p>What really changed wasn&#39;t the CSS, it was where the responsibility lives. With media queries, the responsibility for a component looking good is split between the component and whoever places it: the consumer has to pick the right variant for the slot. With container queries, the responsibility goes back entirely to the component, which is where it belongs.</p>
<p>That&#39;s the difference between a design system that scales and one that turns into a catalog of special cases. Every time a component needs to know where it goes to look right, you&#39;ve created a coupling between the piece and its context. Container queries cut that coupling, and with it, the kind of debt that never shows up in a ticket but slows you down on every new screen.</p>
]]></content:encoded>
    </item>
    <item>
      <title>EventBridge: decoupling services without building a queue monster</title>
      <link>https://yohangel.com/en/blog/eventbridge-desacoplar-servicios/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/eventbridge-desacoplar-servicios/</guid>
      <description>When an event bus beats calling services directly or chaining SQS queues, and the trade-offs I accepted adopting it.</description>
      <pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>AWS</category>
      <content:encoded><![CDATA[<p><strong>For a long time my default way to connect two services was the obvious one: service A calls service B. A POST, a Lambda invoking another, a write that fires a webhook. It works until you have five consumers of the same event and every new consumer forces you to touch the producer. That&#39;s where EventBridge changed my architecture: the producer stops knowing who&#39;s listening. It emits &quot;this happened&quot; to a bus and forgets about it. Consumers subscribe on their own with pattern rules. I chose this model in several AWS services because the coupling between teams was costing me more than the extra latency; in exchange, I accepted that debugging an event-driven system is harder than following a synchronous call. This article is about when it&#39;s worth it and when it isn&#39;t.</strong></p>
<h2>The problem it solves: the producer shouldn&#39;t know its consumers</h2>
<p>Picture a simple domain event: <code>OrderConfirmed</code>. When an order is confirmed, you need to bill, notify the customer, reserve inventory, and update analytics. The naive version has the orders service call all four. The day marketing wants to hear about it too, someone has to open the orders code, add the fifth call, test it, and deploy it. The producer piles up responsibilities that aren&#39;t its own.</p>
<p>With an event bus, the orders service emits a single event and is done. Each consumer decides whether it cares. Adding marketing means creating a new rule pointing at its target, without touching orders. Coupling shifts from &quot;code to code&quot; to &quot;event schema to event schema,&quot; which is far cheaper to maintain.</p>
<h2>The rules pattern: filter on the bus, not in the consumer</h2>
<p>What I liked most about EventBridge over a classic SNS is that filtering lives in the rule, declaratively, over the event content. I don&#39;t hand the event to a Lambda so it can decide whether it matters: the bus only delivers what matches.</p>
<pre><code class="language-json">{
  &quot;source&quot;: [&quot;orders.service&quot;],
  &quot;detail-type&quot;: [&quot;OrderConfirmed&quot;],
  &quot;detail&quot;: {
    &quot;total&quot;: [{ &quot;numeric&quot;: [&quot;&gt;&quot;, 500] }],
    &quot;country&quot;: [&quot;ES&quot;, &quot;MX&quot;, &quot;CO&quot;]
  }
}
</code></pre>
<p>That rule only fires for confirmed orders over 500 in three countries. The consumer doesn&#39;t run even once for the rest. That <code>if</code> used to live inside the function, invoked millions of times to discard most cases; now the discard is free and doesn&#39;t show up in my Lambda bill or my logs.</p>
<h2>Where EventBridge is NOT the answer</h2>
<p>Being honest about the trade-offs is what separates an architecture decision from an act of faith. EventBridge isn&#39;t free on complexity.</p>
<p>If you need the consumer&#39;s response, don&#39;t use a bus. It&#39;s fire-and-forget: you emit and you don&#39;t know what happened next unless you set up another event coming back. For a request/response flow —the user is waiting for a result on screen— a direct synchronous call is still the right thing. Dropping an event bus in there only adds latency and a state machine nobody asked for.</p>
<p>I also don&#39;t use it when strict ordering is a hard requirement. EventBridge doesn&#39;t guarantee order, and it can deliver the same event more than once. If your case needs to process things in exact sequence, an SQS FIFO queue fits better. In fact, a pattern that works for me is EventBridge for the fan-out and SQS as a buffer in front of each slow consumer, combining the best of both.</p>
<blockquote>
<p>💡 The question that decides everything: does the producer need to know what happened to the consumer? If the answer is yes, you don&#39;t have an event, you have a call. Don&#39;t dress it up as event-driven.</p>
</blockquote>
<h2>At-least-once delivery forces you to be idempotent</h2>
<p>This is the detail people discover in production, not in design. EventBridge guarantees at-least-once delivery, which means your consumer can receive <code>OrderConfirmed</code> twice for the same order. If you bill on every reception, you just charged twice.</p>
<p>The solution isn&#39;t praying it won&#39;t happen: it&#39;s designing the consumer so that processing the same event twice yields the same result as processing it once. An idempotency key per event, a table that records &quot;I already processed this ID,&quot; and an <code>INSERT ... ON CONFLICT DO NOTHING</code> in Postgres before doing the work with side effects.</p>
<pre><code class="language-typescript">async function handle(event: OrderConfirmed) {
  const inserted = await db.processedEvents.createIfAbsent(event.id);
  if (!inserted) return; // already processed, exit clean
  await bill(event);
}
</code></pre>
<p>It&#39;s not elegant code, it&#39;s code that survives reality. Any event-driven architecture that doesn&#39;t treat idempotency as a first-class requirement is building on sand.</p>
<h2>Observability: the real price I paid</h2>
<p>The most expensive thing about adopting EventBridge wasn&#39;t the service, it was observability. In a synchronous call you follow the stack trace and see the whole path. On a bus, the producer emitted and left; the consumer failed three hops later; and correlating the two requires that you propagated a <code>correlationId</code> in every event&#39;s <code>detail</code> from day one.</p>
<p>I learned it late and paid for it in incidents that cost extra hours just because I couldn&#39;t join cause to effect. Now every event carries its correlation ID, every consumer logs it, and a DLQ per rule captures failures so no event is silently lost. With that, the decoupling pays off. Without it, you&#39;re trading a coupling problem for a blind-debugging one, and it&#39;s not clear you come out ahead.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Server Actions in Next.js: where they shine and where I refuse to use them</title>
      <link>https://yohangel.com/en/blog/nextjs-server-actions-limites/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/nextjs-server-actions-limites/</guid>
      <description>Server Actions eliminate API boilerplate, but they are not a universal replacement for endpoints. My criteria for deciding.</description>
      <pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>Frontend</category>
      <content:encoded><![CDATA[<p><strong>Next.js Server Actions are the framework&#39;s best and worst idea, depending on where you put them. For form mutations — create, edit, delete from the UI — they eliminate an entire layer of boilerplate: no endpoint to define, no manual serialization, no fetch or loading states managed with useEffect. But I&#39;ve seen (and had to undo) projects where they became the only way to talk to the server, and there the pattern breaks: business logic trapped in the framework, impossible to consume from another client, and an invalidation graph nobody understands. My criteria after using them in production: Server Actions are the UI&#39;s interaction layer, not your product&#39;s API. This article is the line I draw and why.</strong></p>
<h2>Where they shine: the classic form mutation</h2>
<p>The happy case is a form that mutates data and revalidates the view. With an action, shared Zod validation, and React&#39;s own pending state:</p>
<pre><code class="language-typescript">// app/projects/actions.ts
&#39;use server&#39;;

import { revalidatePath } from &#39;next/cache&#39;;
import { projectSchema } from &#39;@/schemas/project&#39;;

export async function createProject(formData: FormData) {
  const session = await getSession();
  if (!session) throw new Error(&#39;Unauthorized&#39;);

  const parsed = projectSchema.safeParse(Object.fromEntries(formData));
  if (!parsed.success) {
    return { error: parsed.error.flatten().fieldErrors };
  }

  await prisma.project.create({
    data: { ...parsed.data, ownerId: session.userId },
  });

  revalidatePath(&#39;/projects&#39;);
  return { ok: true };
}
</code></pre>
<p>Zero endpoints, zero HTTP client, the return type travels inferred all the way to the component. For 80% of the mutations in an internal dashboard, this is all you need, and going back to writing route handlers for those cases is nostalgia, not engineering.</p>
<h2>The red line: business logic living in the action</h2>
<p>The problem starts when the action stops being an adapter and becomes the home of the logic. An action that calculates prices, orchestrates three services, and decides business rules is code that can only be invoked from a React component in that Next.js project. The day the mobile app arrives, or the cron job, or the team that wants a public endpoint, that logic has to be exhumed.</p>
<p>My structural rule: the action calls a service, and the service doesn&#39;t know Next.js exists.</p>
<pre><code class="language-typescript">// services/projects.ts — framework-agnostic, testable in isolation
export async function createProjectForUser(
  input: ProjectInput,
  userId: string
) {
  // all business logic lives here
}
</code></pre>
<p>The action ends up as three lines: authenticate, parse, delegate. If tomorrow I need to expose the same thing via a route handler or an SQS worker, the service is already there. I chose this separation knowing it duplicates a bit of ceremony; in exchange, no business decision is held hostage by the framework.</p>
<blockquote>
<p>💡 A Server Action is a transport mechanism with syntactic sugar, not an architecture layer. If deleting Next.js from your project takes business logic with it, the logic was in the wrong place.</p>
</blockquote>
<h2>The operational limits nobody tells you about until production</h2>
<p><strong>They are sequential POSTs.</strong> Server Actions execute serially by default from the same client: if the user fires three, they queue. For a form it doesn&#39;t matter; for high-frequency interactions (autosave, drag and drop on a board) it&#39;s a bottleneck you discover late.</p>
<p><strong>They are neither cancelable nor cacheable.</strong> A GET to a route handler can be CDN-cached and aborted with AbortController. An action cannot. Anything that&#39;s a read — search, autocomplete, filters — has no business being in an action: that&#39;s a route handler or directly a Server Component reading from the database.</p>
<p><strong>Error handling is opaque.</strong> A throw in an action reaches the client as a generic error in production (by design, to avoid leaking internals). That forces you to model expectable errors as return values — the <code>{ error }</code> in the example — and reserve throw for the truly exceptional. It&#39;s more Result than Exception, and it&#39;s worth deciding on day one so you don&#39;t mix styles.</p>
<h2>The criteria in one sentence</h2>
<p>A mutation initiated by a user from that same project&#39;s UI: Server Action delegating to a service. Reads, APIs consumed by third parties, webhooks, high frequency, anything that might one day need another client: route handler or a proper backend. With that line drawn, Server Actions are an excellent tool — small, boring, and in their place, which is exactly how I like my tools.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Serverless observability: structured logs or debugging blind</title>
      <link>https://yohangel.com/en/blog/observabilidad-serverless-logs/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/observabilidad-serverless-logs/</guid>
      <description>In Lambda there is no server to SSH into. How I structure logs, correlate requests, and find failures without losing my mind.</description>
      <pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>AWS</category>
      <content:encoded><![CDATA[<p><strong>On a traditional server, when something fails, you always have the last resort: SSH, <code>tail -f</code>, and read. In serverless that resort doesn&#39;t exist. A request can cross API Gateway, a Lambda, an SQS queue, and another Lambda — and if your logs are loose strings written with <code>console.log</code>, reconstructing what happened is archaeology. I learned this the worst way: an intermittent production bug that took me days to locate because I couldn&#39;t follow a request end to end. This article is the logging system I&#39;ve used since: structured logs, a correlation ID that travels with the request, and CloudWatch queries that answer questions in minutes.</strong></p>
<h2>Logs as data, not prose</h2>
<p>The fundamental shift is to stop writing logs for humans and start writing them for machines. A prose log (<code>&quot;Error processing order 123 for user 456&quot;</code>) is pleasant to read but impossible to query at scale. A structured log is a JSON object with consistent fields:</p>
<pre><code class="language-typescript">logger.error(&#39;order_processing_failed&#39;, {
  orderId: order.id,
  userId: user.id,
  errorCode: err.code,
  correlationId: ctx.correlationId,
  durationMs: Date.now() - start,
});
</code></pre>
<p>CloudWatch Logs Insights understands JSON natively. With structured logs, &quot;how many orders failed this week with this error code, and from which users?&quot; stops being a heroic grep and becomes a query:</p>
<pre><code class="language-sql">fields @timestamp, orderId, userId
| filter level = &#39;error&#39; and errorCode = &#39;PAYMENT_TIMEOUT&#39;
| stats count(*) by userId
| sort count(*) desc
</code></pre>
<p>I don&#39;t use a homemade logger: AWS Lambda Powertools for TypeScript already solves serialization and levels, and automatically adds the invocation context (request ID, function name, cold start). Reinventing that is wasted time.</p>
<h2>The correlation ID: the thread that ties the system together</h2>
<p>In an event-driven architecture, the problem isn&#39;t logging: it&#39;s correlating. The user&#39;s request triggers a Lambda that enqueues a message in SQS that&#39;s processed by another Lambda that writes to DynamoDB. Four different log groups, four different request IDs, no visible relationship between them.</p>
<p>My rule: the first entry point generates a <code>correlationId</code> (or adopts the one arriving in the <code>X-Correlation-Id</code> header), and that ID travels with everything. In SQS messages it goes inside the message attributes; in service-to-service calls, in a header; in every log, as a mandatory field. Powertools lets you inject it once into the logger and forget about it: every log from that invocation carries it.</p>
<p>With that, reconstructing the complete story of a request is a single Logs Insights query across all the log groups involved. What used to be an afternoon of archaeology is now thirty seconds.</p>
<blockquote>
<p>💡 Observability isn&#39;t added when there&#39;s an incident: it&#39;s designed beforehand. The day something breaks in production, your logs already are what they are. Every field you didn&#39;t log is a question you can&#39;t answer.</p>
</blockquote>
<h2>Metrics without a hammer: EMF</h2>
<p>For business metrics (orders processed, matchings generated, growing queues) I don&#39;t use CloudWatch API calls, which add latency and cost money per request. I use Embedded Metric Format: you write the metric as part of the JSON log with a special format, and CloudWatch extracts it asynchronously. Zero marginal cost on the hot path, and the metrics show up in dashboards and alarms like any other.</p>
<pre><code class="language-typescript">metrics.addMetric(&#39;MatchingCompleted&#39;, MetricUnit.Count, 1);
metrics.addMetric(&#39;MatchingDurationMs&#39;, MetricUnit.Milliseconds, elapsed);
</code></pre>
<p>The discipline I impose on myself: every Lambda emits at least one success metric and one failure metric with a business name, not a technical one. &quot;Errors&quot; tells me nothing at 3 a.m.; &quot;PaymentTimeouts&quot; does.</p>
<h2>What I decided not to do</h2>
<p>I didn&#39;t set up X-Ray everywhere. I tried it, and at my scale full distributed tracing was more noise than signal: most of my flows have two or three hops, and the correlation ID in structured logs covers them more than well enough. I reserve X-Ray for flows with more hops or when I suspect inter-service latencies the logs can&#39;t explain. It&#39;s an honest trade-off: less automatic visibility in exchange for less cost and less configuration to maintain.</p>
<p>I also don&#39;t centralize logs in an external tool. CloudWatch Logs Insights has a mediocre UX and queries billed per GB scanned, but adding a third party means another shipping pipeline, another contract, and another invoice. While the volume allows it, I prefer the pain I know. The day queries become truly slow or expensive, that decision gets revisited with data: GB scanned per month and minutes lost per incident.</p>
<h2>The result</h2>
<p>None of this is glamorous. It&#39;s three habits: JSON instead of prose, an ID that travels with the request, and business metrics embedded in the logs. But the operational difference is enormous: incidents went from &quot;let me see if I can reproduce it&quot; to &quot;give me the correlation ID and I&#39;ll tell you what happened&quot;. In serverless you can&#39;t SSH into the server, so the system had better tell its own story.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Rerankers: the piece my semantic search was missing</title>
      <link>https://yohangel.com/en/blog/rerankers-busqueda-semantica/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/rerankers-busqueda-semantica/</guid>
      <description>Embeddings retrieve reasonable candidates; the reranker decides which ones deserve the top. Here is how I combine them without doubling latency.</description>
      <pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>IA</category>
      <content:encoded><![CDATA[<p><strong>For months I treated embedding search as if it were the whole system: query to vector, cosine similarity against pgvector, top 10, ship it. It worked... until you looked at the ordering of the results closely. The right candidate was almost always there, but in position 4, or 7 — rarely in position 1. The lesson I resisted accepting: embeddings are an excellent coarse filter and a mediocre fine-grained sorter. The reranker is the piece that separates &quot;the relevant results are on the page&quot; from &quot;the best result is first&quot;. This article is how I integrated one into the JXBS matching engine without wrecking latency.</strong></p>
<h2>Why the bi-encoder falls short</h2>
<p>An embedding compresses an entire document into a fixed vector before knowing what you&#39;re going to ask it. That&#39;s its great advantage (you can precompute and index it) and its structural limit: the compression loses the nuances that only matter against a specific query. Two candidate profiles can be equally &quot;close&quot; to a job opening in vector space for different reasons, and cosine similarity can&#39;t tell whether the closeness comes from what&#39;s essential or what&#39;s incidental.</p>
<p>A reranker is a cross-encoder: it receives the query and the document together, and produces a relevance score by looking at the interaction between the two. It&#39;s far more precise at ordering, and far more expensive: you can&#39;t precompute anything, every query-document pair is an inference. That&#39;s why nobody ranks an entire corpus with a cross-encoder. The pattern is retrieve cheap, reorder expensive over few candidates.</p>
<h2>The architecture: retrieve 50, rerank to 10</h2>
<p>My pipeline ended up in two stages. The first is the one I already had: pgvector retrieves the 50 nearest neighbors. The second passes those 50 through the reranker and keeps the best 10 according to the new score.</p>
<pre><code class="language-sql">-- Stage 1: cheap retrieval, deliberately over-retrieving
SELECT id, titulo, contenido
FROM documentos
ORDER BY embedding &lt;=&gt; $1
LIMIT 50;
</code></pre>
<pre><code class="language-typescript">// Stage 2: reranking over the candidates
const scored = await rerank({
  query: userQuery,
  documents: candidates.map((c) =&gt; c.contenido),
});

const top = scored
  .sort((a, b) =&gt; b.relevanceScore - a.relevanceScore)
  .slice(0, 10);
</code></pre>
<p>The number 50 isn&#39;t sacred. It&#39;s the balance I found between two failure modes: retrieve too few and the right document never even reaches the reranker, so there&#39;s nothing to reorder; retrieve too many and you pay latency and inference for candidates that never had a chance. I chose 50 because in my tests the correct result was inside the bi-encoder&#39;s top 50 practically always; in exchange, I accept that if retrieval fails at the root, the reranker won&#39;t rescue it. Garbage in, garbage out — just better sorted.</p>
<h2>Latency: the real price and how I contained it</h2>
<p>Reranking adds a network call and an inference per request. In my case that meant going from ~80ms to ~350ms at p50. For an interactive search that was too much, and I contained it with three decisions:</p>
<p>First, only rerank when it matters. If the bi-encoder&#39;s top candidate score has a wide lead over the second, the ordering is already clear and I skip the second stage. Second, truncate documents: the reranker doesn&#39;t need the document&#39;s 3,000 words — the title and the first relevant fragment are enough to discriminate, and cost grows with tokens. Third, cache normalized query-document pairs; in a real product, queries repeat far more than you&#39;d think.</p>
<blockquote>
<p>💡 The reranker doesn&#39;t improve your retrieval: it improves your <em>precision at the top</em>. If your problem is that good results don&#39;t even show up in the top 50, your problem is in the embeddings or the chunking, and no reranker will save you.</p>
</blockquote>
<h2>How I measured it was worth it</h2>
<p>Before integrating anything I built a small evaluation set: ~100 real queries with the result a human considered correct marked by hand. The metric I cared about was MRR (mean reciprocal rank): if the correct result is first it scores 1, second scores 0.5, and so on. With the bi-encoder alone I had a mediocre MRR; with the reranker it went up clearly and consistently. I&#39;m not giving exact figures because they depend brutally on the domain and the dataset, but the pattern repeats on almost any corpus: the bi-encoder finds, the cross-encoder orders.</p>
<p>What does generalize is the method: don&#39;t integrate a reranker because it&#39;s fashionable. Build the 100 human-judged queries first, measure your baseline, and let the number make the decision. In my case the number was conclusive and the extra latency was paid for.</p>
<h2>Trade-offs I accepted</h2>
<p>I chose a managed reranker via API instead of serving my own cross-encoder, because at my scale the per-request cost was lower than the fixed cost of keeping a GPU warm. In exchange, I accepted an external dependency in the search&#39;s critical path, mitigated with a fallback: if the reranker doesn&#39;t respond within 500ms, I return the bi-encoder&#39;s ordering as-is. Slightly worse search is infinitely better than search that&#39;s down.</p>
<p>I also accepted that the system is harder to reason about: there are now two models that can degrade independently. That&#39;s why the evals run against the full pipeline, not against each piece in isolation. What reaches the user is the output of the whole chain, and that&#39;s the only thing worth measuring.</p>
]]></content:encoded>
    </item>
    <item>
      <title>File uploads with presigned URLs: let S3 do the heavy lifting</title>
      <link>https://yohangel.com/en/blog/s3-presigned-urls-subidas/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/s3-presigned-urls-subidas/</guid>
      <description>Why I stopped piping files through my backend and how to design the presigned URL flow without security holes.</description>
      <pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>AWS</category>
      <content:encoded><![CDATA[<p><strong>Piping a file through your backend to upload it to S3 is paying the same toll twice. The file travels from the browser to your API, consumes your server&#39;s memory and bandwidth, and then travels again from your API to S3. On a Lambda, you also hit the payload limit; on Fargate, you&#39;re occupying a container to act as a pipe. The alternative I&#39;ve used for years: the backend signs a temporary URL with surgical permissions, the browser uploads directly to S3, and your API only processes metadata. I chose this pattern in the dashboards I build for clients because it scales without touching infrastructure; in exchange, the flow has more steps and the security has to be designed carefully, because a poorly scoped presigned URL is an open door. This article is the full flow.</strong></p>
<h2>The flow in three steps</h2>
<p>The pattern has three actors: the browser, your API, and S3. The API never sees the file.</p>
<pre><code class="language-typescript">// 1. The backend generates the signed URL (Lambda or a regular endpoint)
import { S3Client, PutObjectCommand } from &#39;@aws-sdk/client-s3&#39;;
import { getSignedUrl } from &#39;@aws-sdk/s3-request-presigner&#39;;

const s3 = new S3Client({});

export async function createUploadUrl(userId: string, contentType: string) {
  if (!ALLOWED_TYPES.has(contentType)) {
    throw new Error(&#39;File type not allowed&#39;);
  }

  const key = `uploads/${userId}/${crypto.randomUUID()}`;

  const url = await getSignedUrl(
    s3,
    new PutObjectCommand({
      Bucket: process.env.UPLOADS_BUCKET,
      Key: key,
      ContentType: contentType,
      ContentLength: undefined, // limit via POST policy conditions if needed
    }),
    { expiresIn: 300 } // 5 minutes, not one more
  );

  return { url, key };
}
</code></pre>
<p>The browser does a direct <code>PUT</code> to that URL with the file as the body. When it finishes, it notifies your API with the <code>key</code>, and there you register the file in PostgreSQL via Prisma, trigger async processing, or whatever applies.</p>
<h2>The three security decisions that matter</h2>
<p><strong>The key is generated by the server, always.</strong> If you let the client pick the file name, you&#39;re letting it pick where it writes in your bucket. The key includes the authenticated <code>userId</code> and a UUID: the client contributes not a single character.</p>
<p><strong>Short expiration and pinned content-type.</strong> Five minutes is enough to start any upload. And signing the <code>ContentType</code> into the URL forces what&#39;s uploaded to match what was declared — it doesn&#39;t stop someone from renaming an executable to <code>.jpg</code>, but it closes the easy path.</p>
<p><strong>The uploads bucket is not the final bucket.</strong> Everything lands in a staging bucket with a lifecycle rule that deletes objects after 24 hours. A later process — in my case a Lambda triggered by the <code>s3:ObjectCreated</code> event — validates the file for real (magic bytes, size, scanning if applicable) and moves it to the definitive bucket. Whatever nobody validates self-destructs on its own.</p>
<blockquote>
<p>💡 A presigned URL is not &quot;access to S3&quot;: it&#39;s ONE operation, on ONE key, during ONE interval. If your signed URL allows more than that, you haven&#39;t understood the pattern — you&#39;ve opened a hole with extra steps.</p>
</blockquote>
<h2>The confirmation event: don&#39;t trust the client</h2>
<p>The classic mistake with this pattern is marking the file as &quot;uploaded&quot; when the browser says it finished. The browser lies: the tab closes, the network fails, the user cancels. The source of truth is S3, not the client.</p>
<p>That&#39;s why the database record has two phases: the API creates the row in <code>pending</code> state when signing the URL, and the <code>ObjectCreated</code> event Lambda moves it to <code>available</code> when the object actually exists. Rows stuck in <code>pending</code> for more than an hour get cleaned up by a job. The client can notify to speed up the UX, but its notification is an optimization, never the truth.</p>
<h2>When not to use this pattern</h2>
<p>As always, there are trade-offs. If files are small (JSON, avatars of a few KB) and volume is low, going through the backend is simpler and simplicity wins. If you need to transform the file inline before storing it (resizing, synchronous transcoding), the backend has to see it anyway. And if your frontend is an offline-first PWA, direct upload complicates the sync queue: there I sometimes prefer to queue the file locally and upload it from a service worker when there&#39;s network, which changes the whole design.</p>
<p>For everything else — documents, images, anything that weighs megabytes and arrives with concurrency — letting S3 do the heavy lifting is one of the few architecture decisions I&#39;ve never regretted.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Tool calling in production: the LLM doesn't execute, it proposes</title>
      <link>https://yohangel.com/en/blog/tool-calling-agentes-produccion/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/tool-calling-agentes-produccion/</guid>
      <description>The mental model that avoids half the bugs in a tool-using agent: the LLM suggests calls, your code decides whether they run.</description>
      <pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>IA</category>
      <content:encoded><![CDATA[<p><strong>The most expensive mistake I made putting tool calling into a product was mental, not code: I thought the LLM executed the tools. It doesn&#39;t. The model returns a JSON object saying &quot;I&#39;d like to call <code>searchCandidates</code> with these arguments,&quot; and there its job ends. Who decides whether that call runs, with what permissions, and what happens to the result, is your code. Internalizing that boundary —the LLM proposes, your backend disposes— completely changed how I design agents. I stopped treating the model as an autonomous executor and started treating it as a planner that emits intentions I validate. This article is that mental model and the production decisions that follow from it.</strong></p>
<h2>The boundary that changes everything</h2>
<p>When you give a model tools, the real loop is: you pass it the user&#39;s request and the schema of available tools; the model responds either with text, or with one or more <em>tool calls</em> (name + arguments in JSON); your code executes whichever ones it decides to; you return the result to the model; and the model continues. The LLM never touches your database or your API. It&#39;s a generator of structured intentions.</p>
<p>That boundary is the best news of the architecture, because it means all the security control lives on your side. The model can <em>ask</em> to delete a record; whether it gets deleted depends on an <code>if</code> of yours. Treating the tool call as a request subject to authorization, and not as an order, is what separates an agent you can put in production from a demo that scares you.</p>
<h2>Schemas are your contract, and the model reads them</h2>
<p>The quality of a tool-using agent depends brutally on how well described the tools are. The model chooses what to call based on the name, the description, and the parameter schema. Vague descriptions produce wrong calls; precise descriptions produce correct calls.</p>
<pre><code class="language-typescript">const tools = [{
  name: &#39;search_candidates&#39;,
  description: &#39;Search candidates by skills and seniority. &#39; +
    &#39;Use it only when the user asks for specific profiles, &#39; +
    &#39;not for general questions about the market.&#39;,
  input_schema: {
    type: &#39;object&#39;,
    properties: {
      skills: { type: &#39;array&#39;, items: { type: &#39;string&#39; } },
      seniority: { type: &#39;string&#39;, enum: [&#39;junior&#39;, &#39;mid&#39;, &#39;senior&#39;] },
    },
    required: [&#39;skills&#39;],
  },
}];
</code></pre>
<p>That <code>enum</code> isn&#39;t decoration: it restricts the output space and makes it far less likely that the model invents a value your backend can&#39;t handle. Every restriction you put in the schema is a class of bug you eliminate before it happens. I&#39;ve learned to invest in strict schemas with the same seriousness I design a public API, because to the model that&#39;s exactly what they are.</p>
<h2>Validate the output as if it came from a hostile client</h2>
<p>Here&#39;s the point most people skip. The model generates the arguments as text, and even if you give it a schema, there&#39;s no hard guarantee it respects your business invariants. It can hand you a valid <code>seniority</code> per the enum but an empty <code>skills</code> array, or a correctly formatted date that&#39;s in the past.</p>
<p>That&#39;s why I validate every tool call with Zod before executing it, with the same distrust I&#39;d apply to user input. The tool call is user input, just generated by a model. If it doesn&#39;t validate, I don&#39;t execute: I return the error to the model as the tool result and let it retry with corrected arguments. That &quot;you failed validation, here&#39;s why, try again&quot; loop is surprisingly effective and keeps the system safe without human intervention.</p>
<blockquote>
<p>💡 A tool call is an untrusted request that happens to come from an LLM instead of a browser. Validate it, authorize it, and log it exactly like any external input. The model is not part of your trust boundary.</p>
</blockquote>
<h2>Side effects: separate reading from writing</h2>
<p>Not all tools are equal, and treating them as if they were is asking for trouble. I distinguish two categories with different rules. Read-only ones —search, query, calculate— I run without friction: if the model gets the query wrong, the cost is a poor answer, nothing irreversible. Write ones —create, update, delete, send— go through a much stricter filter.</p>
<p>For write tools with real impact, the pattern that works for me is not to give the model the destructive tool directly, but one that <em>proposes</em> the action and leaves it pending an explicit confirmation. The model can draft the email; sending it requires a step my code controls, often with a human in the loop. I chose this caution in exchange for less &quot;magical&quot; agents, and I&#39;d choose it again: an agent that can send messages without supervision is an incident waiting its turn.</p>
<h2>Why this mental model scales and the other doesn&#39;t</h2>
<p>When you think the LLM executes, every new tool scares you, because you&#39;re widening what a non-deterministic system can do on its own. When you understand that the LLM only proposes, adding tools is cheap: each one is one more intention your code knows how to validate, authorize, and execute under your rules.</p>
<p>The agent stops being a black box you pray to and becomes a planner whose proposals pass through a control layer you wrote and understand. All of the model&#39;s creativity, none of its ability to cause harm without permission. That division of responsibilities isn&#39;t an implementation detail: it&#39;s the difference between an agent you can defend in a security review and one you can&#39;t.</p>
]]></content:encoded>
    </item>
    <item>
      <title>One Zod schema to rule them all: shared client/server validation</title>
      <link>https://yohangel.com/en/blog/zod-validacion-compartida/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/zod-validacion-compartida/</guid>
      <description>Duplicating validation between the form and the API is a source of silent bugs. Here is how I share a single schema across both worlds.</description>
      <pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>Frontend</category>
      <content:encoded><![CDATA[<p><strong>Every serious form validates twice: on the client, for immediate feedback, and on the server, because the client is hostile territory and you never trust it. The problem is when those two validations are two different implementations. In a dashboard I built for a client, the frontend accepted a phone number with spaces and the backend rejected it: the user saw the field turn green and the submit failed with a cryptic error. It wasn&#39;t a code bug, it was a duplication bug. Since then I follow one rule: the validation schema is written ONCE, in Zod, in a shared package, and both the form and the API import it. This article is that pattern.</strong></p>
<h2>The real problem: two sources of truth that diverge</h2>
<p>Duplicated validation doesn&#39;t fail the day you write it: it fails six months later, when someone changes a rule in the backend (a character minimum, a new postal code format) and nobody touches the frontend. From then on you have two behaviors: what the form says is valid and what the API actually accepts. The user suffers the difference.</p>
<p>The solution isn&#39;t discipline (&quot;remember to change both places&quot;), because discipline doesn&#39;t scale. The solution is structural: make it so only one place exists.</p>
<h2>The schema as a shared contract</h2>
<p>In a monorepo, the schema lives in its own package, with no React or Node dependencies:</p>
<pre><code class="language-typescript">// packages/schemas/src/registro.ts
import { z } from &#39;zod&#39;;

export const registroSchema = z.object({
  email: z.string().email(&#39;Invalid email&#39;),
  telefono: z
    .string()
    .transform((v) =&gt; v.replace(/\s+/g, &#39;&#39;))
    .pipe(z.string().regex(/^\+?\d{9,15}$/, &#39;Invalid phone number&#39;)),
  password: z.string().min(12, &#39;Minimum 12 characters&#39;),
});

export type RegistroInput = z.input&lt;typeof registroSchema&gt;;
export type RegistroOutput = z.output&lt;typeof registroSchema&gt;;
</code></pre>
<p>Two details matter here. First, the <code>transform</code> + <code>pipe</code>: normalization (stripping spaces from the phone) is part of the schema, so the &quot;phone with spaces&quot; is valid in both worlds and gets normalized the same way in both worlds. Second, exporting <code>z.input</code> and <code>z.output</code> as separate types: what goes in (with spaces) and what comes out (normalized) are different types, and TypeScript forces you not to confuse them.</p>
<h2>On the client: the same schema drives the form</h2>
<p>With react-hook-form and its Zod resolver, the form validates against the shared schema without writing a single extra rule:</p>
<pre><code class="language-typescript">const form = useForm&lt;RegistroInput&gt;({
  resolver: zodResolver(registroSchema),
});
</code></pre>
<p>The error messages you defined in the schema appear under each field. If tomorrow the password minimum goes up to 14, you change one number in one file and the form and the API find out at the same time.</p>
<h2>On the server: parse, don&#39;t trust</h2>
<p>In the API, the same schema is the boundary between the outside world and your typed code:</p>
<pre><code class="language-typescript">export async function POST(req: Request) {
  const parsed = registroSchema.safeParse(await req.json());
  if (!parsed.success) {
    return Response.json(
      { errors: parsed.error.flatten().fieldErrors },
      { status: 422 },
    );
  }
  // parsed.data is RegistroOutput: normalized and typed
  await crearUsuario(parsed.data);
}
</code></pre>
<p>I use <code>safeParse</code> and return <code>fieldErrors</code> with the same structure the form consumes. That way, even if a malicious client skips browser validation, the error the server returns fits into the same per-field error UI. Client validation is UX; server validation is security. Same logic, different roles.</p>
<blockquote>
<p>💡 &quot;Parse, don&#39;t validate&quot;: the schema doesn&#39;t just check that the data is correct, it <em>transforms</em> it into a type that can no longer be incorrect. Past the boundary, no function ever asks again whether the email is valid: the type guarantees it.</p>
</blockquote>
<h2>The limits of the pattern</h2>
<p>Not everything is shareable, and pretending it is produces monstrous schemas. Validations that depend on server state (does this email already exist? is this coupon still active?) don&#39;t belong in the shared schema: they&#39;re backend business rules and get checked there, returning errors in the same <code>fieldErrors</code> format so the UI renders them identically.</p>
<p>I also decided not to share API response schemas with the same enthusiasm. Input schemas change slowly and you control them; typing every response with Zod at runtime adds a parsing cost that buys nothing on most endpoints. There, the TypeScript type contract generated from the schema itself is enough for me.</p>
<p>The pattern&#39;s overall trade-off is coupling: frontend and backend share a package, which demands a monorepo or publishing a versioned package. I chose a monorepo with pnpm workspaces because the cost of coordinating versions across separate repos was exactly the kind of friction I was trying to eliminate. In exchange, a schema change forces a coordinated deploy of both sides. I accept it: I prefer an explicit coordinated deploy over a silent divergence discovered by a user.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Islands in Astro: when to hydrate and when not to</title>
      <link>https://yohangel.com/en/blog/astro-islands-hidratacion/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/astro-islands-hidratacion/</guid>
      <description>Astro's partial hydration isn't magic: it's a per-component decision you make, and getting it wrong costs JavaScript nobody asked for.</description>
      <pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>Frontend</category>
      <content:encoded><![CDATA[<p><strong>Astro starts from a premise that, after years of SPAs, sounds almost heretical: by default, your component ships no JavaScript to the browser. None. It renders to HTML at build time and stays there, static and fast. Interactivity isn&#39;t the page&#39;s natural state, it&#39;s something you request explicitly, component by component, with a <code>client:</code> directive. This blog is built that way, and understanding the islands architecture — when to hydrate and, above all, when not to — is the difference between a page that flies and one that drags kilobytes of framework around just to animate a button.</strong></p>
<h2>The mental model: HTML by default, JS on demand</h2>
<p>In a classic SPA everything is JavaScript: the framework mounts entirely on the client, hydrates the whole tree, and governs the page from there. It&#39;s comfortable for the developer and expensive for the user, who pays the cost of booting the framework even when 90% of the screen is text that never changes.</p>
<p>Astro flips the premise. It renders your components to HTML on the server or at build time, and by default discards the JavaScript that generated them. What reaches the browser is plain HTML that paints instantly. The &quot;islands&quot; are the specific zones that do need a life of their own — a carousel, a form with validation, a dropdown menu — surrounded by a sea of static HTML that costs nothing. Each island hydrates independently; the rest of the page never notices.</p>
<h2>The <code>client:</code> directives are the decision, not a detail</h2>
<p>The lever is the directive you give the component when you use it. Each one describes <em>when</em> that island hydrates, and choosing well is where the performance lives:</p>
<pre><code class="language-astro">---
import Carousel from &#39;../components/Carousel.jsx&#39;;
import NewsletterForm from &#39;../components/NewsletterForm.jsx&#39;;
import HeavyChart from &#39;../components/HeavyChart.jsx&#39;;
---

&lt;!-- Hydrates as soon as the page loads: for what&#39;s above the fold --&gt;
&lt;Carousel client:load /&gt;

&lt;!-- Hydrates when it enters the viewport: for what&#39;s further down --&gt;
&lt;HeavyChart client:visible /&gt;

&lt;!-- Hydrates when the browser is idle: for non-urgent stuff --&gt;
&lt;NewsletterForm client:idle /&gt;
</code></pre>
<p><code>client:load</code> hydrates immediately: use it only for what&#39;s interactive from the first visible pixel. <code>client:idle</code> waits for the main thread to breathe. <code>client:visible</code> is my favorite: it doesn&#39;t spend a byte of JavaScript until the component enters the screen, ideal for everything below the fold. The directive isn&#39;t a syntax ornament: it&#39;s literally the load policy of each island, and placing them carelessly is how JavaScript the user doesn&#39;t need sneaks in.</p>
<h2>When NOT to hydrate (which is almost always)</h2>
<p>The most common mistake for people arriving from React is treating every component as if it needs state. It doesn&#39;t. A card, a heading, a list of posts, a footer: all of that is HTML that paints once and never changes again. If there&#39;s no event, there&#39;s no state and no interaction, there&#39;s no island. It stays as a static Astro component and ships zero JavaScript.</p>
<p>The question I ask of every component isn&#39;t &quot;can it be interactive?&quot; but &quot;does it have to respond to something from the user on the client?&quot;. A button that&#39;s just a link isn&#39;t an island, it&#39;s an <code>&lt;a&gt;</code> tag. An accordion can be built with native <code>&lt;details&gt;</code> and <code>&lt;summary&gt;</code> without a line of JS. The more I push interactivity toward native HTML and CSS, the fewer islands I need and the lighter the page stays.</p>
<blockquote>
<p>💡 In Astro, every <code>client:</code> you write is JavaScript the user downloads. The best optimization isn&#39;t hydrating faster, it&#39;s not hydrating.</p>
</blockquote>
<h2>Islands don&#39;t share state, and that&#39;s on purpose</h2>
<p>Here&#39;s the trade-off you have to be clear on before marrying the model. Each island is an isolated component: it hydrates on its own and doesn&#39;t magically share state with the others. If I&#39;m coming from a SPA where a global store wires everything together, this feels like a limitation. Two islands that need to talk to each other don&#39;t do it through props, because they&#39;re independent trees mounted separately.</p>
<p>The solution exists — browser events, shared <code>nanostores</code>, signals — but it&#39;s real friction, and that friction is the sign that maybe you&#39;re using the tool against the grain. If your page is basically an app with tightly interwoven state, many parts talking to each other constantly, Astro will ask you to do gymnastics for something a SPA framework gives you for free. There I don&#39;t force Astro: I pick the right tool.</p>
<h2>Why I chose this model anyway</h2>
<p>Astro shines when content leads and interactivity is occasional: blogs, marketing sites, documentation, portfolios, e-commerce with product islands. That&#39;s exactly the shape of most of the web, and of this site. I accept the friction of coordinating islands in exchange for every page being born weightless by default, and for me being the one who decides and justifies every kilobyte of JavaScript I add.</p>
<p>That shift of the default is everything. In a SPA, JavaScript is the norm and removing it takes work; in Astro, JavaScript is the exception and adding it is a conscious decision. I&#39;d rather have a system where the cheap thing is the easy thing and the expensive thing requires me to ask for it on purpose. That inversion of the default burden is, for the kind of sites I build, the best performance decision I can make without optimizing a single line.</p>
]]></content:encoded>
    </item>
    <item>
      <title>View Transitions in Astro: SPA-like navigation without the SPA weight</title>
      <link>https://yohangel.com/en/blog/astro-view-transitions/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/astro-view-transitions/</guid>
      <description>How I get smooth page-to-page transitions and state that survives navigation without giving up the HTML Astro serves by default.</description>
      <pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>Frontend</category>
      <content:encoded><![CDATA[<p><strong>I chose Astro for several sites precisely because it serves HTML and ships very little JavaScript to the client. The classic downside of that model is that every click reloads the whole page: white flash, jumping scroll, lost state. Astro&#39;s View Transitions gave me the good part of a SPA —fluid navigation, no visible reloads, state that persists— without dragging back a client router and megabytes of bundle. This article is how I use them and, above all, where the edges are that you&#39;ll want to know before turning them on.</strong></p>
<h2>What problem they actually solve</h2>
<p>In a traditional MPA site every navigation is a full document reload. It works, it&#39;s robust, and it&#39;s what Astro does out of the box, but it feels rough: between one page and the next there&#39;s a blank instant, the browser resets the scroll, and anything playing gets cut off. In a SPA that doesn&#39;t happen because you never reload: a router intercepts the click and replaces the content with JavaScript. The price is that you take on the whole SPA apparatus —the router, hydration, client state— even if your site is mostly content.</p>
<p>View Transitions sit in the middle. Astro intercepts the navigation, fetches the next page, swaps the <code>&lt;body&gt;</code> content without reloading the document, and animates the transition using the browser&#39;s View Transitions API. You keep serving normal HTML pages, but navigating between them stops flashing. Turning it on is one line in the layout.</p>
<pre><code class="language-astro">---
// src/layouts/Layout.astro
import { ClientRouter } from &#39;astro:transitions&#39;;
---
&lt;html lang=&quot;en&quot;&gt;
  &lt;head&gt;
    &lt;ClientRouter /&gt;
  &lt;/head&gt;
  &lt;body&gt;
    &lt;slot /&gt;
  &lt;/body&gt;
&lt;/html&gt;
</code></pre>
<p>With that <code>&lt;ClientRouter /&gt;</code> in the <code>&lt;head&gt;</code>, every internal link now navigates without a reload and with a smooth default transition. You haven&#39;t written a router or turned anything into a SPA: it&#39;s still Astro serving HTML.</p>
<h2>Named animations: the detail that sells it</h2>
<p>The default transition is a discreet fade, but what hooks people is being able to animate specific elements from one page to another. If I mark two elements on different pages with the same <code>transition:name</code>, the browser understands they&#39;re &quot;the same&quot; and animates the change in position and size between them. The typical case is an article thumbnail that grows into the detail page&#39;s header.</p>
<pre><code class="language-astro">&lt;!-- In the listing --&gt;
&lt;img src={post.cover} transition:name={`cover-${post.slug}`} /&gt;

&lt;!-- On the detail page --&gt;
&lt;img src={post.cover} transition:name={`cover-${post.slug}`} /&gt;
</code></pre>
<p>The name has to be unique per element on each page; that&#39;s why I derive it from the <code>slug</code>. If you repeat the same <code>transition:name</code> for several elements visible at once, the browser can&#39;t tell which is which and the animation breaks. It&#39;s the most common mistake starting out.</p>
<h2>The state you want to survive</h2>
<p>Here&#39;s the less obvious win. Since there&#39;s no full reload, I can ask certain elements to persist across navigations instead of being recreated. An audio player, a heavy map, a menu with its open state: with <code>transition:persist</code> the element is kept as-is when moving between pages, without resetting.</p>
<pre><code class="language-astro">&lt;audio src={song} controls transition:persist /&gt;
</code></pre>
<p>Without this, every navigation would destroy the <code>&lt;audio&gt;</code> and recreate it, cutting the playback. With <code>transition:persist</code>, the same DOM node travels to the next page and keeps playing. It&#39;s exactly the kind of continuity people build a whole SPA for, and here I get it with an attribute.</p>
<blockquote>
<p>💡 Every persisted element is a node you decide not to recreate. Use it for things that need real continuity —audio, video, an expensive map— and not as a patch to avoid recomputing something that should actually refresh with the new page.</p>
</blockquote>
<h2>The edges you have to know</h2>
<p>None of this is free, and it&#39;s fair to say where it pinches. The first is scripting. Since the page doesn&#39;t reload, the <code>&lt;script&gt;</code> that ran on <code>DOMContentLoaded</code> doesn&#39;t run again on each navigation: it ran once and the new page doesn&#39;t fire it again. Any initialization you took for granted on every load has to be re-hooked to the <code>astro:page-load</code> event, which Astro emits on the initial load and after each transition.</p>
<pre><code class="language-typescript">// Runs on first load and after every navigation
document.addEventListener(&#39;astro:page-load&#39;, () =&gt; {
  initWidgets();
});
</code></pre>
<p>This is the bug that&#39;s hardest to spot, because the site works perfectly on a manual reload and only breaks when navigating between pages: the scripts simply weren&#39;t re-hooked.</p>
<p>The second edge is degradation. View Transitions rely on an API that not every browser implements the same way; where it&#39;s missing, Astro gracefully falls back to a normal navigation with reload. That means you can&#39;t treat the transition as guaranteed: it&#39;s a progressive enhancement, not a base to build logic on. And the third, more judgment than technical, is not overdoing it: animating everything is dizzying. I reserve named animations for one or two visual connections that genuinely add continuity, and leave the rest as the discreet fade.</p>
<h2>Why it&#39;s worth it to me</h2>
<p>I chose View Transitions in Astro rather than jumping to a SPA because I keep what I wanted from the framework —served HTML, minimal JavaScript, independent pages that cache and index well— and on top of that gain the fluidity that used to require changing the entire model. In exchange I accept a handful of concrete edges: re-hooking scripts on <code>astro:page-load</code>, treating the transition as progressive, and being disciplined with animations. It&#39;s a trade I almost always accept, because the cost is bounded and known, whereas moving into a SPA just to get an animation would mean paying in bundle, complexity, and SEO a price far above the problem it solves.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Queues with SQS: retries that don't duplicate work</title>
      <link>https://yohangel.com/en/blog/colas-sqs-idempotencia/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/colas-sqs-idempotencia/</guid>
      <description>How I make SQS retries safe in JXBS without sending duplicate emails, charges, or state changes.</description>
      <pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>AWS</category>
      <content:encoded><![CDATA[<p><strong>A queue isn&#39;t a luxury: it&#39;s how you pull everything that doesn&#39;t need an instant answer off the critical path. In JXBS I send to SQS whatever can wait a few seconds: candidate emails, match recomputation, syncs. What I learned the hard way is that the tricky part isn&#39;t enqueuing, it&#39;s making sure the retry doesn&#39;t redo the work. SQS delivers &quot;at least once&quot;, and that &quot;at least&quot; is exactly where things break.</strong></p>
<h2>Why I enqueue instead of answering inline</h2>
<p>When a recruiter posts a role, the HTTP response shouldn&#39;t wait for embeddings to be generated, matches recomputed, and notifications fired. I send that to a queue and return immediately. The user perceives a fast app; the heavy lifting happens behind the scenes, at its own pace.</p>
<p>The queue gives me three concrete things. It absorbs spikes: if a hundred roles come in at once, they&#39;re processed at the rate the consumer can handle, not all at once. It isolates failures: if the email provider is down, the task retries itself without taking down the original request. And it decouples: the producer neither knows nor cares who consumes.</p>
<h2>The real problem: &quot;at least once&quot;</h2>
<p>Standard SQS guarantees delivery, not uniqueness. The same message can arrive twice: because the consumer took longer than the visibility timeout and SQS redelivered, because of a network retry, or because the process died right after working but before deleting the message.</p>
<p>If my consumer just sends an email per message, a candidate gets two. If it creates a charge, I bill twice. Duplicate delivery isn&#39;t a rare edge case that happens once a year: it&#39;s the system&#39;s normal behavior, and you have to design assuming it will happen.</p>
<h2>Idempotency: the only defense that scales</h2>
<p>The fix isn&#39;t to avoid duplicates, it&#39;s to make processing twice yield the same result as processing once. That&#39;s idempotency, and I implement it with a stable key per unit of work.</p>
<pre><code class="language-sql">CREATE TABLE processed_message (
  idempotency_key TEXT PRIMARY KEY,
  processed_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);
</code></pre>
<p>Before doing any work, the consumer tries to insert the key. If it already exists, the message is a duplicate and I discard it with no side effects.</p>
<pre><code class="language-typescript">async function handle(msg: Job) {
  const inserted = await db.processedMessage
    .create({ data: { idempotencyKey: msg.id } })
    .catch(() =&gt; null); // hits the PK if already processed

  if (!inserted) return; // duplicate: do nothing

  await doTheActualWork(msg);
}
</code></pre>
<p>I don&#39;t invent the key at random: I derive it from the message&#39;s intent. For a welcome email it&#39;s <code>welcome-email:{candidateId}</code>, not a fresh UUID per attempt. That way two messages representing the same action collide and only one wins.</p>
<blockquote>
<p>💡 A safe retry isn&#39;t one that never repeats; it&#39;s one that can repeat without anyone caring.</p>
</blockquote>
<h2>Visibility timeout and the DLQ</h2>
<p>Two SQS pieces need tuning. The visibility timeout must be longer than the realistic worst-case processing time: if I take longer, SQS assumes I failed and redelivers while I&#39;m still working. I size it against the consumer&#39;s p99, not the average.</p>
<p>And the dead-letter queue: after N failed attempts, the message leaves the main queue and lands in a DLQ. Without it, a poison message —one that always fails— retries forever and clogs the queue. The DLQ is where I inspect what broke without it blocking everything else. I put an alarm on it: if something lands there, I want to know.</p>
<h2>What I&#39;d do differently</h2>
<p>At first I put the idempotency logic inside each handler. I ended up repeating it and getting it wrong differently in each one. Today it lives in a single wrapper: it takes the key, checks, executes, and no handler thinks about duplicates again. The pattern matters more than the infrastructure. SQS is replaceable; the discipline of assuming every message can arrive twice is what actually saves you.</p>
]]></content:encoded>
    </item>
    <item>
      <title>DynamoDB single-table: the pattern that took me longest to get</title>
      <link>https://yohangel.com/en/blog/dynamodb-single-table/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/dynamodb-single-table/</guid>
      <description>Why cramming several entities into one DynamoDB table stops looking insane once you think in access patterns instead of entities.</description>
      <pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>AWS</category>
      <content:encoded><![CDATA[<p><strong>I come from PostgreSQL, so the first time I saw a DynamoDB single-table design — users, orders, and products living in the same table, with keys like <code>USER#123</code> and <code>ORDER#456</code> — I thought someone had lost their mind. It took me a while because I was asking the wrong question: in a relational model you design by entities, and in DynamoDB you design by access patterns. Once that clicks, the single-table stops being an aberration and becomes the natural way to get value out of the database. This article is how I made that mental shift and where the traps are.</strong></p>
<h2>Why DynamoDB isn&#39;t Postgres with different skin</h2>
<p>In Postgres I normalize, create tables per entity, and let the planner resolve joins at query time. It&#39;s flexible: if tomorrow I need a new query, I write it and the engine figures it out, maybe with an extra index. I pay for that flexibility with performance that depends on the plan the engine picks.</p>
<p>DynamoDB flips the deal. There are no joins and no planner: every access is a lookup by key, and if you didn&#39;t design the key for that query, that query is either impossible or forces a <code>Scan</code> that walks the whole table. In return, a well-designed access has predictable latency no matter how much data there is. The consequence is brutal: in DynamoDB you can&#39;t start from the data model, you have to start from the list of queries your product is going to make.</p>
<h2>Start from the access patterns, always</h2>
<p>Before touching a table, I write the list of accesses in plain language. For an orders product it&#39;d be something like: give me a user by id; give me all of a user&#39;s orders; give me an order with its line items; give me a user&#39;s orders sorted by date. That list isn&#39;t documentation, it&#39;s the design. Each pattern has to map to a partition-key <code>Query</code>, not a <code>Scan</code>.</p>
<p>This step is the one people coming from SQL skip, and it&#39;s exactly the one you can&#39;t skip. If a pattern you didn&#39;t foresee shows up later, in Postgres it&#39;s a new query; in DynamoDB it can be a key redesign or a global secondary index. The cost of getting this wrong is high, so this is where I spend the thinking.</p>
<h2>Generic keys: PK and SK that don&#39;t mean one single thing</h2>
<p>The single-table trick is that the partition key and sort key aren&#39;t called <code>user_id</code> or <code>order_id</code>. They&#39;re called plain <code>PK</code> and <code>SK</code>, and their contents change depending on the entity. A user lives at <code>PK = USER#123</code>, <code>SK = PROFILE</code>. Their orders live in the same partition, <code>PK = USER#123</code>, <code>SK = ORDER#456</code>. So asking for a user and all their orders is a single <code>Query</code> on <code>PK = USER#123</code>: the partition already brings the profile and the orders together, no join.</p>
<pre><code class="language-typescript">// All of a user&#39;s items (profile + orders) in one query
const res = await ddb.query({
  TableName: &#39;app&#39;,
  KeyConditionExpression: &#39;PK = :pk&#39;,
  ExpressionAttributeValues: { &#39;:pk&#39;: &#39;USER#123&#39; },
});

// Orders only: narrow by the sort key prefix
const orders = await ddb.query({
  TableName: &#39;app&#39;,
  KeyConditionExpression: &#39;PK = :pk AND begins_with(SK, :prefix)&#39;,
  ExpressionAttributeValues: { &#39;:pk&#39;: &#39;USER#123&#39;, &#39;:prefix&#39;: &#39;ORDER#&#39; },
});
</code></pre>
<p>The <code>begins_with</code> on the sort key is the lever: by prefixing the types (<code>ORDER#</code>, <code>PROFILE</code>, <code>ADDRESS#</code>) I can pull exactly the subset I want from a partition, sorted, in a single call.</p>
<blockquote>
<p>💡 In DynamoDB the sort key doesn&#39;t sort data, it designs queries. The prefix you give it is what decides which questions you&#39;ll be able to ask cheaply.</p>
</blockquote>
<h2>GSIs: when you need to look at the data along another axis</h2>
<p>Patterns that don&#39;t fit the primary key are solved with global secondary indexes. A GSI is essentially a projection of the table with a different PK and SK, maintained by DynamoDB asynchronously. If I need &quot;all orders with status <code>PENDING</code> sorted by date,&quot; I create a GSI whose PK is the status and whose SK is the date, and that access is a clean <code>Query</code> again.</p>
<p>The pattern I use is index overloading: the same generic <code>GSI1PK</code> and <code>GSI1SK</code> attributes mean different things depending on the entity, just like the primary key. With a couple of well-thought-out GSIs I cover almost any product. What I don&#39;t do is create a GSI on a whim: each index costs writes and storage, because every <code>put</code> on the table replicates to the indexes that apply.</p>
<h2>The trade-offs nobody tells you about up front</h2>
<p>Single-table is powerful, but it&#39;s fair to admit what it costs. The first thing is the mental curve: modeling this way is uncomfortable until you internalize thinking in accesses, and a new team takes a while to read a table where everything is called <code>PK</code> and <code>SK</code>. The second is rigidity: if a truly unforeseen access pattern appears, adapting costs more than in SQL, where an ad-hoc query is always possible even if it&#39;s slow.</p>
<p>That&#39;s why I don&#39;t use single-table for everything. When a product has exploratory queries, shifting reports, or relationships I can&#39;t anticipate, Postgres gives me a flexibility that DynamoDB would charge dearly for. I pick DynamoDB when the access patterns are known, bounded, and I need predictable latency at scale without operating a relational engine. There the single-table shines: fewer pieces, flat latency, and a bill that scales with real usage and not with the size of the table. Like everything in AWS, it isn&#39;t the best tool, it&#39;s the best tool for a problem with the right shape.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Lambda or Fargate: how I decide which to use</title>
      <link>https://yohangel.com/en/blog/lambda-vs-fargate/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/lambda-vs-fargate/</guid>
      <description>I stopped asking which is better and started asking what shape the workload has. The answer almost always falls out of that.</description>
      <pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>AWS</category>
      <content:encoded><![CDATA[<p><strong>&quot;Lambda or Fargate?&quot; is one of those questions people ask expecting a winner, and there isn&#39;t one. Both run code without you managing servers, but they solve differently shaped problems. After building plenty of services on ECS Fargate and plenty of functions on Lambda, I stopped choosing by fashion and started looking at three things about the workload: how traffic arrives, how long each unit of work lasts, and how much state the process needs to hold. Almost always the decision falls out on its own.</strong></p>
<h2>It isn&#39;t serverless versus containers</h2>
<p>The first misunderstanding is treating it as serverless versus containers. Fargate is also serverless: you define a task, AWS runs it, and you never touch an EC2 instance. The real difference is the execution model. Lambda is ephemeral and event-invoked: a request arrives, an environment spins up, your handler runs, it responds, and it shuts down. Fargate is a long-lived process: your container starts once, sits listening, and serves requests until you scale it or kill it.</p>
<p>That difference colors everything. In Lambda you don&#39;t control the lifecycle: there&#39;s no &quot;startup&quot; of yours where you open a connection pool and calmly reuse it, because the environment can be recycled at any moment. In Fargate you do have a stable process with its startup, its warm memory, and its connection pool alive between requests. Choosing is, at bottom, choosing how much control you want over that lifecycle.</p>
<h2>The shape of traffic rules</h2>
<p>The first thing I look at is how the work arrives. If traffic is intermittent, bursty, or unpredictable —a webhook that fires now and then, a cron that processes on schedule, an endpoint with odd spikes— Lambda fits naturally. It scales from zero to many parallel invocations without you configuring anything, and when there&#39;s no traffic you pay nothing. That &quot;zero to a thousand without warning&quot; is exactly where Fargate struggles: holding capacity for the spike means paying for containers sitting idle, and reactive scaling is slower than a Lambda cold start.</p>
<p>If traffic is sustained and constant, the math flips. A service under steady load all day, on Lambda, is a continuous drip of invocations billed by the millisecond that ends up expensive next to a handful of Fargate containers at full tilt. There a long-lived process squeezes more out of each CPU: it doesn&#39;t repeat the startup cost on every request and keeps warm the resources Lambda would rebuild over and over.</p>
<blockquote>
<p>💡 Lambda charges per invocation and duration: it shines when traffic is sometimes zero. Fargate charges per running container: it shines when traffic is almost never zero. Before deciding, draw a day&#39;s traffic curve.</p>
</blockquote>
<h2>Duration and limits: the 15-minute clock</h2>
<p>The second axis is how long each unit of work lasts. Lambda has a hard cap of fifteen minutes per invocation. For an API or event processing that&#39;s plenty, but any job that might run over —a long migration, video processing, a job that walks a large dataset— doesn&#39;t fit, and slicing it to fit is complexity that doesn&#39;t always pay off. Fargate has no such clock: a task can run for minutes or hours, which makes it the natural home for long batch work or processes that simply have to stay alive.</p>
<p>There&#39;s also the cold start. When Lambda spins up a fresh environment, the first invocation pays the cost of initializing the runtime. For latency-tolerant loads it doesn&#39;t matter; for an endpoint sensitive to tail latency, that occasional cold hit shows. It&#39;s mitigated —provisioned concurrency, lightweight runtimes— but it&#39;s a real factor. Fargate has no per-request cold start: you pay startup once when you scale a new task, not on every invocation.</p>
<h2>State, connections, and dependencies</h2>
<p>The third axis is how much state and which dependencies the process needs. The classic case is the database. A connection pool lives on reusing connections between requests, and that fits a long-lived process like Fargate. In Lambda, each concurrent environment opens its own, and under a spike you can exhaust PostgreSQL&#39;s connections without noticing; the fix is an intermediate connection proxy, one more piece in the diagram. It&#39;s not impossible to run Lambda against a relational database, but it&#39;s friction Fargate doesn&#39;t have.</p>
<pre><code class="language-typescript">// In Fargate this initializes ONCE at container start
// and the pool is reused on every request.
import { Pool } from &#39;pg&#39;;

const pool = new Pool({ max: 10 }); // lives between requests

export async function handler(req: Request) {
  const client = await pool.connect(); // reuses, doesn&#39;t reopen
  try {
    return await query(client, req);
  } finally {
    client.release();
  }
}
</code></pre>
<p>That same pattern in Lambda is treacherous: the <code>pool</code> survives while the environment is reused, but it multiplies per concurrent environment, and that&#39;s where the database blows up under load.</p>
<h2>My rule, and why I mix the two</h2>
<p>In the end my criterion is simple. I start on Lambda when the work is short, event-triggered, with irregular traffic and little state: webhooks, async tasks, glue between AWS services. I move to Fargate when the work is sustained, long-running, or needs live connections and lifecycle control: APIs under constant load, queue workers processing nonstop, long batch.</p>
<p>And I don&#39;t pick one for the whole system. In the same product an HTTP service on Fargate with a stable Postgres connection coexists with Lambda functions for incoming webhooks and async jobs. I chose to mix rather than force a single model because each workload has its shape, and fighting the tool to do something it isn&#39;t built for always costs you: in Lambda with proxies and slicing, in Fargate with idle capacity waiting for a spike that never comes. Like almost everything in AWS, there&#39;s no best tool, there&#39;s the best tool for the shape of your problem.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Evals for LLM features: how I know I didn't break anything</title>
      <link>https://yohangel.com/en/blog/llm-evals-produccion/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/llm-evals-produccion/</guid>
      <description>Without a set of evals, every prompt or model change is a blind bet. Here is how I built mine.</description>
      <pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>IA</category>
      <content:encoded><![CDATA[<p><strong>The first time I changed a prompt in production at JXBS, I did it like I&#39;d change anything else: tried it against two or three hand-picked examples, decided it looked better, and shipped. A week later a user reported that matching returned garbage on a case that used to work. I hadn&#39;t broken the code: I&#39;d broken the model&#39;s behavior, and I had no way of knowing because I wasn&#39;t measuring anything. Evals are the piece that turns &quot;working with LLMs&quot; from superstitious craft into engineering. This article is the minimum viable set of evals I use today on every feature that has a model inside it.</strong></p>
<h2>An LLM has no unit tests, and that&#39;s the problem</h2>
<p>When you write a normal function, a unit test pins down its contract: this goes in, that comes out, and if someone breaks it the pipeline warns you. With an LLM-based feature that contract blurs. The same prompt with the same input can give different answers, and a seemingly harmless change —reordering two sentences of the system prompt, bumping from one model version to the next— can degrade cases you didn&#39;t even know depended on that detail.</p>
<p>The mental mistake is treating the prompt as stable code and the model as a fixed dependency. They&#39;re neither. The prompt is a parameter you tune constantly, and the model is a dependency that shifts under you without warning. Without a battery of cases you run on every change, you&#39;re shipping blind and discovering regressions through your users. Which is the worst way to discover them.</p>
<h2>The minimum eval: a golden dataset and a metric you care about</h2>
<p>You don&#39;t need an eval platform to start. You need a file with representative cases and a function that scores. A case is a real input, the expected output (or a property the output must satisfy) and optionally why it&#39;s there. I pull them from three places: examples I know work, edge cases that worry me, and above all the real bugs people report —every production regression becomes an eval case so it can&#39;t happen again.</p>
<pre><code class="language-typescript">type EvalCase = {
  name: string;
  input: MatchInput;
  // An assertion about the output, not exact equality:
  // with LLMs, comparing strings word by word is useless.
  expect: (output: MatchOutput) =&gt; boolean;
};

const cases: EvalCase[] = [
  {
    name: &#39;senior candidate with exact stack ranks first&#39;,
    input: seniorReactExact,
    expect: (out) =&gt; out.ranked[0].id === &#39;cand_42&#39;,
  },
  {
    name: &#39;does not invent skills absent from the CV&#39;,
    input: cvWithoutKubernetes,
    expect: (out) =&gt; !out.ranked[0].reasons.includes(&#39;Kubernetes&#39;),
  },
];
</code></pre>
<p>Notice I don&#39;t compare the whole output character by character. With an LLM that&#39;s doomed to fail over irrelevant wording differences. What I check are properties: that the right candidate lands on top, that no hallucinated skill shows up, that the format is parseable. Each assertion encodes something I genuinely care about in the behavior, not the exact shape of the text.</p>
<h2>Scoring the subjective: when I use an LLM as a judge</h2>
<p>Many outputs can&#39;t be verified with an <code>if</code>. &quot;Is this explanation of the match clear and well-grounded?&quot; is not an obvious boolean property. For that I use a second LLM as an evaluator: I pass it the input, the output and a concrete rubric, and ask for a score with justification. It works surprisingly well for catching large degradations, and it&#39;s far cheaper than reviewing hundreds of outputs by hand.</p>
<p>But I chose to use LLM-as-judge with my eyes open, not as a silver bullet. It has known biases: it tends to reward long answers, favors style over substance, and if you ask it to score its own model it turns lenient. I treat it as a noisy filter, not the truth: it&#39;s good for catching big quality drops between versions, not for telling whether something went from 8.2 to 8.4. For the properties I can verify with code, I use code, which is deterministic and free.</p>
<blockquote>
<p>💡 Every LLM bug that reaches production is an eval case you were missing. Before fixing the prompt, add the failing case to your dataset. That way the fix is locked in and that specific regression can never come back.</p>
</blockquote>
<h2>Running evals where it hurts: before you deploy</h2>
<p>An eval dataset you run when you remember is worthless. The value shows up when it runs automatically on every change that touches the prompt, the model or the orchestration logic. I wired it into the pipeline as one more step: if the pass rate over the dataset falls below a threshold, the deploy stops just as it would if a test failed.</p>
<pre><code class="language-typescript">async function runEvals(cases: EvalCase[]) {
  const results = await Promise.all(
    cases.map(async (c) =&gt; ({
      name: c.name,
      passed: c.expect(await runFeature(c.input)),
    })),
  );

  const passed = results.filter((r) =&gt; r.passed).length;
  const rate = passed / results.length;

  console.table(results);
  if (rate &lt; 0.9) {
    throw new Error(`Evals below threshold: ${rate}`);
  }
}
</code></pre>
<p>The threshold isn&#39;t 100% on purpose. With LLMs there are genuinely ambiguous cases where an occasional miss is tolerable, and demanding a perfect green pushes you to overfit the prompt to your dataset. I prefer a high but realistic threshold and I watch the trend: if the rate has been dropping for three versions, there&#39;s a problem even if each version cleared the bar.</p>
<h2>Where I put the effort today</h2>
<p>If I had to give a single piece of advice to someone shipping their first LLM into a product, it would be this: build the evals before polishing the prompt, not after. It&#39;s tempting to spend the time on the creative part of writing clever instructions, but without a way to measure you&#39;re optimizing blind. I chose to invest in the evaluation harness before the perfect prompt because the harness pays for itself on the first regression it catches, and its value grows with every case you add. I&#39;ll rewrite the prompt ten times; the evals tell me which of those ten versions is better.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Structured outputs with LLMs: how I stopped parsing text by hand</title>
      <link>https://yohangel.com/en/blog/llm-structured-outputs/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/llm-structured-outputs/</guid>
      <description>Why function calling and Zod validation turned my LLMs-in-product from a fragile component into a reliable piece of infrastructure.</description>
      <pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>IA</category>
      <content:encoded><![CDATA[<p><strong>The first time I put an LLM into a product, the model did its job well and I did mine badly: I asked for JSON in the prompt, prayed, and then wrote regex to rescue the object from between apologies and markdown blocks. It worked 90% of the time, which in production is another way of saying it failed. The fix wasn&#39;t a smarter prompt — it was treating the model&#39;s output as a typed interface: structured outputs with function calling and Zod validation. This article is how I made that shift and what I got for it.</strong></p>
<h2>The real problem isn&#39;t the model, it&#39;s the edge</h2>
<p>An LLM inside a product doesn&#39;t live alone: its output feeds a function, a query, or a call to another API. That contact point between natural language and typed code is the fragile edge. If the model returns free text, the edge fills up with defensive parsing: find the first <code>{</code>, count braces, strip the ```json, catch the <code>JSON.parse</code> that blows up. Every one of those lines is debt that breaks the moment the model decides to get conversational.</p>
<p>For a while I tried to solve it with prompt engineering. &quot;Respond with valid JSON only, no explanations.&quot; It helps, but it guarantees nothing: the prompt is a plea, not a contract. The lesson that took me a while to accept is that reliability isn&#39;t requested, it&#39;s enforced at the API layer.</p>
<h2>Function calling as a contract, not a trick</h2>
<p>Modern models accept a tool or response schema and commit to returning something that fits it. That changes the game: instead of describing the format in prose, I declare it as structure and the model fills in the blanks. I stop parsing intentions and start receiving data.</p>
<p>I model it like this: I define the shape I need, pass it as a response schema, and the rest of my code assumes it will receive exactly that. The prompt goes back to talking about the task — what to extract, how to prioritize — and not about commas and quotes.</p>
<pre><code class="language-typescript">import { z } from &#39;zod&#39;;

const CandidateExtraction = z.object({
  seniority: z.enum([&#39;junior&#39;, &#39;mid&#39;, &#39;senior&#39;, &#39;lead&#39;]),
  primary_stack: z.array(z.string()).max(8),
  years_experience: z.number().int().min(0).max(50),
  location: z.string(),
  open_to_remote: z.boolean(),
});

type CandidateExtraction = z.infer&lt;typeof CandidateExtraction&gt;;
</code></pre>
<p>That schema is the source of truth. From it comes the TypeScript type, from it comes the JSON schema I hand the model, and against it I validate the response. One definition, three uses.</p>
<h2>Zod: the net that catches what the model doesn&#39;t guarantee</h2>
<p>Here&#39;s the nuance a lot of people skip: the API accepting a schema doesn&#39;t mean the result is correct in the sense I care about. The model can return structurally valid JSON that&#39;s semantically absurd: a <code>years_experience</code> of 200, a stack with fifteen duplicate entries, an enum that never existed. The shape is right; the content isn&#39;t.</p>
<p>That&#39;s why I always validate at the edge with Zod, even when using structured outputs. It isn&#39;t belt-and-suspenders paranoia: it&#39;s that the API&#39;s guarantee and my domain&#39;s guarantee are different things. Zod lets me express the second one — ranges, max lengths, closed enums — and turns a dubious response into an error I can handle before it pollutes the database.</p>
<pre><code class="language-typescript">const raw = await callModelWithSchema(prompt, CandidateExtraction);
const parsed = CandidateExtraction.safeParse(raw);

if (!parsed.success) {
  // I don&#39;t propagate garbage: I retry with the error as context,
  // or fall back to a manual flow. I never write without validating.
  return handleInvalid(parsed.error);
}

await saveCandidate(parsed.data); // parsed.data is already the correct type
</code></pre>
<blockquote>
<p>💡 An LLM in a product isn&#39;t a source of truth, it&#39;s a source of proposals. Validation is what turns a proposal into data you trust.</p>
</blockquote>
<h2>Retries with the error as context</h2>
<p>When validation fails, the instinct is to retry with the same prompt. You waste tokens repeating the same mistake. What actually works is re-injecting the failure: I hand the model back what I expected and what it gave me, and ask it to fix it. In practice, 90% of validation failures resolve in a single informed retry, because they&#39;re almost always dumb slips — a missing field, a misspelled enum — that the model corrects immediately once you point at the exact spot.</p>
<p>I cap the retries, of course. If after two attempts it still doesn&#39;t fit, the problem isn&#39;t the model having a bad day: it&#39;s that the task is poorly framed or the schema asks for something the input doesn&#39;t contain. There, the loud failure is a gift, because it forces me to fix the cause instead of hiding it.</p>
<h2>The trade-off I accepted</h2>
<p>Structured outputs isn&#39;t free. Boxing the model into a schema reduces its flexibility: for genuinely open-ended tasks, where I don&#39;t know the shape of the answer ahead of time, the schema gets in the way more than it helps. And adding Zod plus retries is code and latency I didn&#39;t have before.</p>
<p>I chose structured outputs anyway because in a product I almost never want creativity at the edge: I want determinism. I&#39;d rather pay in rigidity and a few lines of validation in exchange for the rest of my system being able to treat the LLM&#39;s output the way it treats any other typed API. That&#39;s the real goal: the model stops being a special exception everyone tiptoes around and becomes just another component, with its contract and its error handling, like any other piece I already trust.</p>
]]></content:encoded>
    </item>
    <item>
      <title>OpenTofu modules you won't hate six months from now</title>
      <link>https://yohangel.com/en/blog/opentofu-modulos-iac/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/opentofu-modulos-iac/</guid>
      <description>Most infrastructure modules age badly. These are the limits I set for myself so they don't.</description>
      <pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>AWS</category>
      <content:encoded><![CDATA[<p><strong>An infrastructure module is one of those things you write once and suffer many times. You start with a clean module to stand up a service on ECS Fargate, and six months later it&#39;s a monster with forty variables, three booleans that toggle incompatible behaviors, and a nested <code>count</code> nobody understands. I migrated several projects from Terraform to OpenTofu and along the way rewrote almost all my modules, and what I learned isn&#39;t about syntax: it&#39;s about resisting the temptation to make one module do everything. This is the set of rules I follow today so infrastructure as code stays an asset and not a millstone.</strong></p>
<h2>A module models a concept, not a list of resources</h2>
<p>The mistake I made most often was building modules around &quot;things I usually deploy together&quot; instead of around a concept with a clear boundary. A &quot;web-service&quot; module that creates the ECS service, its load balancer, its DNS, its database and its alarms feels convenient, but it couples decisions that change at different rates. The day you want a service without a database, or one sharing a database with another, the module fights you.</p>
<p>The rule that works for me is that a module should model a concept worth naming on its own. &quot;A service on Fargate&quot; is a concept. &quot;An SQS queue with its DLQ&quot; is a concept. &quot;Everything project X needs&quot; is not: it&#39;s a composition, and compositions belong in the root, not inside a module. When each module has a clean conceptual boundary, composing them is easy and swapping one for another doesn&#39;t break the rest.</p>
<h2>Configuration booleans are debt in disguise</h2>
<p>Every time you add an <code>enable_something</code> variable to a module, you&#39;re introducing a fork in its behavior. With two or three flags it&#39;s still understandable. With eight, your module has hundreds of possible combinations and you&#39;ve only tested the four you use. The rest are minefields waiting for someone.</p>
<pre><code class="language-hcl"># The module that ages badly: configured with flags
# that toggle incompatible branches of behavior.
module &quot;service&quot; {
  source          = &quot;./modules/service&quot;
  enable_https    = true
  enable_autoscaling = true
  enable_spot     = false
  enable_efs      = true
  # ...and on until nobody knows which combinations work
}
</code></pre>
<p>I prefer smaller modules with fewer options, and I resolve variation by composing instead of configuring. If one service needs persistent storage and another doesn&#39;t, I don&#39;t add an <code>enable_efs</code> flag: I make the volume a separate module I wire in when needed. The service module knows nothing about EFS and has no untested branch. The composition shows up in the root, where it should, not hidden behind a boolean.</p>
<h2>Outputs are the contract; treat them like an API</h2>
<p>What a module exposes in its <code>outputs</code> is its public interface, and changing it breaks consumers just as you&#39;d break the clients of an API. Early on I exposed everything &quot;just in case&quot;, and ended up with modules whose consumers depended on internal details I wanted to change. Now I treat outputs with the same discipline as a contract: I expose the minimum a consumer needs to wire this module to another —an ARN, an id, an endpoint— and none of the internal plumbing.</p>
<pre><code class="language-hcl"># Expose identifiers to compose with, not whole resources.
output &quot;service_arn&quot; {
  value = aws_ecs_service.this.id
}

output &quot;task_role_arn&quot; {
  # Other modules attach policies to this role;
  # that&#39;s the extension point I want to offer.
  value = aws_iam_role.task.arn
}
</code></pre>
<p>Exposing a role&#39;s ARN instead of the whole role is a small but important example: it gives the consumer the exact point to extend (attach a policy) without opening the full resource for them to poke at things that would break my module.</p>
<blockquote>
<p>💡 Before adding a variable to a module, ask whether the variation belongs inside the module or in whatever composes it. Most of the time it belongs outside. A module with few inputs and clear outputs gets reused; one with twenty flags gets copy-pasted.</p>
</blockquote>
<h2>State is where design mistakes get paid</h2>
<p>In IaC the design isn&#39;t truly tested until you apply changes over existing infrastructure, and that&#39;s where state sends you the bill. A badly bounded module shows itself when a <code>plan</code> wants to destroy and recreate something you only meant to nudge, because you moved a resource around or swapped a <code>count</code> for a <code>for_each</code>. Migrating to OpenTofu taught me to treat state structure as part of the design, not a detail: <code>moved</code> blocks exist precisely to refactor without destroying, and using them carefully is what lets you reorganize modules without a scare in production.</p>
<p>I chose small, clearly bounded modules knowing the price: there are more pieces to compose in the root and composition is more verbose than one mega-module with flags. In exchange, each piece is understandable, testable and replaceable on its own, and a change in one doesn&#39;t threaten to recreate half the infrastructure. For code that will live for years and that I apply changes to every week, that predictability is worth far more than the convenience of writing fewer lines on day one.</p>
<h2>Where I put the effort today</h2>
<p>If I had to sum it all up in one sentence: design modules around how they&#39;ll be composed and changed, not around how much you type today. Infrastructure as code isn&#39;t a script you run once, it&#39;s a codebase you maintain, and it suffers the same ills as any other —coupling, untested branches, contracts that break— with the added twist that here a mistake can take down a service. Fewer options, clear boundaries and outputs treated like an API are what keep a module yours six months from now instead of being the corner of the repo nobody wants to touch.</p>
]]></content:encoded>
    </item>
    <item>
      <title>RAG in production: chunking matters more than the model</title>
      <link>https://yohangel.com/en/blog/rag-chunking-produccion/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/rag-chunking-produccion/</guid>
      <description>How you split your documents decides more about RAG quality than the LLM you put on top.</description>
      <pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>IA</category>
      <content:encoded><![CDATA[<p><strong>When I built the first serious RAG for semantic search in JXBS, I spent days comparing embedding models and trying ever-bigger LLMs on top of the retriever. The real quality jump didn&#39;t come from there. It came from fixing how I split documents before indexing them. Chunking —that boring part everyone solves with a character <code>split</code> and moves on from— is the decision that most determines whether your RAG answers well or hallucinates with confidence. This article is what I learned by retrieving useless fragments the hard way.</strong></p>
<h2>The retriever can only return what you indexed well</h2>
<p>A RAG has two halves: retrieve the relevant fragments and generate an answer from them. The second half gets all the attention because that&#39;s where the flashy LLM lives. But generation is bounded by what reaches it: if the retriever brings back three mediocre fragments, no model, however large, will invent the missing context. At best it will fake it, which is worse.</p>
<p>And what the retriever can return depends entirely on how you cut the text. An embedding represents the meaning of a whole fragment. If that fragment mixes two distinct ideas, its vector lands at an in-between point that represents neither well, and similarity search skips it exactly when you need it. The problem isn&#39;t the embedding model: it&#39;s that you handed it a piece of text with no clear idea inside.</p>
<h2>Splitting by characters is the default mistake</h2>
<p>The recipe almost every tutorial ships is: cut every 1000 characters with 200 of overlap. It&#39;s convenient and it&#39;s what I did at first. The problem is that 1000 characters mean nothing semantically: the cut falls mid-sentence, separates a definition from its example, or puts the end of one section and the start of the next in the same chunk. You end up with vectors representing fragments split down the middle.</p>
<p>The first cheap fix is to split by structure, not by length. Documents already come with semantic boundaries: paragraphs, headings, list items. Cutting along those boundaries makes each chunk hold a reasonably complete idea.</p>
<pre><code class="language-typescript">// Instead of blindly chopping every N characters,
// respect the document&#39;s natural boundaries
function chunkByStructure(markdown: string): string[] {
  // Split by section headings first
  const sections = markdown.split(/\n(?=#{1,3}\s)/);

  return sections.flatMap((section) =&gt; {
    // A short section is a whole chunk
    if (section.length &lt;= 1200) return [section];
    // A long one subdivides by paragraphs, not characters
    return section
      .split(/\n\n+/)
      .reduce&lt;string[]&gt;((acc, para) =&gt; {
        const last = acc[acc.length - 1];
        if (last &amp;&amp; (last + &#39;\n\n&#39; + para).length &lt;= 1200) {
          acc[acc.length - 1] = last + &#39;\n\n&#39; + para;
        } else {
          acc.push(para);
        }
        return acc;
      }, []);
  });
}
</code></pre>
<p>It isn&#39;t sophisticated, but the quality jump over blind cutting is immediate because each vector now represents a unit of meaning rather than an arbitrary slice.</p>
<h2>Chunk size is a trade-off, not a magic number</h2>
<p>There&#39;s no correct value here, there&#39;s a tension you have to resolve for your case. Small chunks give very precise embeddings —the vector represents a concrete idea— but they fragment context: the answer to a question can be spread across five slices and the retriever only brings back the three most similar. Large chunks preserve context but dilute the embedding: the more text you cram in, the more the meaning averages out and the less the search discriminates.</p>
<p>My practical rule is to start with medium chunks, the size of a couple of paragraphs covering a single subtopic, and adjust by watching what the system retrieves on real queries. If I see correct answers getting cut off, I raise the size; if I see fragments that talk about several things at once, I lower it. Overlap between adjacent chunks helps so an idea straddling two slices isn&#39;t lost, but overlap doesn&#39;t fix a cut made in the wrong place.</p>
<blockquote>
<p>💡 Don&#39;t optimize chunk size in the abstract. Write ten real queries from your product, look at which fragments the system retrieves for each, and adjust from there. A manual eval of ten cases tells you more than any heuristic.</p>
</blockquote>
<h2>Metadata and context: the chunk doesn&#39;t live alone</h2>
<p>A fragment ripped from the middle of a document loses where it came from. &quot;The term is thirty days&quot; means nothing without knowing the term of what. That&#39;s why I attach metadata to every chunk —document title, section, date— and where I can I prepend a line of context to the text that gets embedded, so the vector also captures what the fragment belongs to. That small header makes ambiguous fragments retrievable by the right query.</p>
<p>Metadata also gives you pre-filtering. In pgvector I can combine similarity search with a <code>WHERE</code> over regular columns, narrowing the search space before comparing vectors.</p>
<pre><code class="language-sql">SELECT id, content
FROM chunks
WHERE document_type = &#39;contract&#39;
  AND language = &#39;en&#39;
ORDER BY embedding &lt;=&gt; $1
LIMIT 5;
</code></pre>
<p>Filtering by metadata before ordering by distance doesn&#39;t just improve precision: it shrinks the set the vector search runs over, and with it the latency.</p>
<h2>Where I put the effort today</h2>
<p>If I had to divide up the time again, I&#39;d spend the bulk of it on the ingestion phase —how I split, what metadata I store, how I enrich each chunk with its context— and far less trying models. The embedding model matters, and the generator LLM matters, but both operate on what ingestion prepares for them. I chose to invest in chunking instead of a bigger model because chunking is cheap to iterate and its effect is a multiplier: it improves every query the system serves, today and in the future. In exchange, it&#39;s unglamorous work that&#39;s hard to show off. It&#39;s still worth it to me.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Modeling the domain with types: when the compiler writes your tests</title>
      <link>https://yohangel.com/en/blog/typescript-tipos-dominio/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/typescript-tipos-dominio/</guid>
      <description>A good type model makes impossible states fail to compile. Here is how I use TypeScript for that.</description>
      <pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>Frontend</category>
      <content:encoded><![CDATA[<p><strong>Most of the TypeScript I see uses types as an autocomplete layer: an <code>interface</code> here, an <code>any</code> there when it gets in the way, and off we go. It works, but it wastes the most valuable thing the language offers. If you model the domain well, the compiler stops being a spellchecker and becomes the thing that prevents states that shouldn&#39;t exist from existing. In the dashboards and PWAs I build, a whole class of bugs I used to catch with tests or in review now simply won&#39;t compile. This article is how I came to use the type system as part of the design, not as decoration.</strong></p>
<h2>Impossible states shouldn&#39;t be representable</h2>
<p>The pattern that most changed how I write code is this: if a state can&#39;t happen in your domain, make it so it can&#39;t be written in your type either. The canonical example is the classic &quot;loading / loaded / error&quot;. The naive version models it with loose optional fields:</p>
<pre><code class="language-typescript">// The type allows states that make no sense:
// loading true and data present at once, or error with data.
type State = {
  loading: boolean;
  data?: Result;
  error?: Error;
};
</code></pre>
<p>That type admits <code>{ loading: true, data: something, error: somethingElse }</code>, a state that should never exist but which the compiler happily accepts. Every component that consumes it has to remember to check the combinations by hand, and the day someone forgets, a bug appears. The discriminated union closes the door:</p>
<pre><code class="language-typescript">// Now each variant carries exactly the data it has,
// and no impossible combination is representable.
type State =
  | { status: &#39;loading&#39; }
  | { status: &#39;success&#39;; data: Result }
  | { status: &#39;error&#39;; error: Error };
</code></pre>
<p>With this version, accessing <code>data</code> without first checking that <code>status</code> is <code>&#39;success&#39;</code> won&#39;t compile. The compiler forces you to handle each case, and the &quot;loading with data and error at the same time&quot; state simply can&#39;t be written. I haven&#39;t added a single test and I&#39;ve eliminated an entire family of bugs.</p>
<h2>Branded types: not all strings are equal</h2>
<p>A <code>string</code> can be a user id, an email or a token, and to TypeScript they&#39;re interchangeable. That means passing an id where an email was expected compiles without complaint, and that error is found at runtime or never. Branded types put a label on the type so the compiler distinguishes things that are structurally equal but conceptually not.</p>
<pre><code class="language-typescript">type UserId = string &amp; { readonly __brand: &#39;UserId&#39; };
type Email = string &amp; { readonly __brand: &#39;Email&#39; };

function sendInvitation(to: Email) { /* ... */ }

const id = &#39;usr_123&#39; as UserId;
sendInvitation(id); // ❌ won&#39;t compile: UserId is not Email
</code></pre>
<p>The trick is that <code>__brand</code> doesn&#39;t exist at runtime —it&#39;s pure type, zero cost— but it forces a <code>UserId</code> and an <code>Email</code> not to be confused. Combined with validation at the boundary (a function that validates a string and returns a branded <code>Email</code>), you get the type system to guarantee, from that point on, that whatever flows around is already validated. Validation stops being something you &quot;hope someone did earlier&quot;.</p>
<blockquote>
<p>💡 When you&#39;re torn between validating at runtime and trusting a type, do both once at the boundary: validate the incoming data and return it branded. From there the compiler propagates that guarantee for free through all your code, without a single extra check.</p>
</blockquote>
<h2>Letting inference work instead of annotating everything</h2>
<p>A common mistake is annotating types everywhere out of habit, even where TypeScript infers them better than you do. Over-annotating isn&#39;t safer: it couples your code to type names that are hard to change later, and sometimes it hides behind a wide type what inference would have kept precise. I prefer to annotate the boundaries —the public signatures of functions, the contracts between modules— and let inference do the work inside.</p>
<pre><code class="language-typescript">// `as const` makes the inferred type exact,
// not a generic `string[]`.
const TAGS = [&#39;IA&#39;, &#39;AWS&#39;, &#39;Frontend&#39;] as const;
type Tag = (typeof TAGS)[number]; // &#39;IA&#39; | &#39;AWS&#39; | &#39;Frontend&#39;
</code></pre>
<p>That pattern lets me keep a single source of truth —the <code>TAGS</code> array— and derive the type from it. If tomorrow I add a tag to the array, the type updates itself and every place doing an exhaustive <code>switch</code> over <code>Tag</code> starts throwing a compile error until I handle the new case. The data and the type can&#39;t drift apart because one is derived from the other.</p>
<h2>The cost, because there always is one</h2>
<p>Modeling the domain with rich types isn&#39;t free. Discriminated unions force you to write more explicit branches, branded types add ceremony at the boundaries, and there&#39;s a point where insisting on expressing an invariant in the type system produces signatures nobody wants to read. I chose this style accepting that cost because in a dashboard I maintain for years the compiler is the only reviewer that never tires or gets distracted, and every invariant I manage to express as a type is a test I don&#39;t have to write or maintain. But I know when to stop: if a type becomes harder to understand than the bug it prevents, that&#39;s where a runtime check and an honest comment win.</p>
<h2>Where I put the effort today</h2>
<p>My rule of thumb is to invest in types where the domain has rules that genuinely matter —which states are valid, which data is validated, which values are interchangeable and which aren&#39;t— and not spend effort typing every last internal detail that inference already covers. Used that way, TypeScript stops being a nuisance asking for annotations and becomes a design tool: you write the model once, and the compiler makes sure that nobody, including you six months from now, can misuse it without noticing.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Cursor + Claude on a real team: integrating AI without losing quality</title>
      <link>https://yohangel.com/en/blog/cursor-claude-equipo-real/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/cursor-claude-equipo-real/</guid>
      <description>What to automate, what to always review, and how to keep the team sharp when AI writes half the code.</description>
      <pubDate>Fri, 12 Jun 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>IA</category>
      <content:encoded><![CDATA[<p><strong>When we adopted Cursor and Claude on the team, the question was never whether AI could write code — we already knew it could. The real question was different: how do we stop a team that produces twice the code from producing twice the bugs?</strong></p>
<p>After months working this way in a monorepo with thousands of commits — where a meaningful share of the code goes through AI-assisted workflows — here is what worked, what failed, and what I would do differently from day one.</p>
<h2>AI doesn&#39;t replace judgment, it demands it</h2>
<p>The first mistake we made was treating AI like just another junior developer: hand it a task and trust the result. The code compiled, the tests passed… and it still broke project conventions that were written down nowhere, because they lived in the team&#39;s heads.</p>
<p>The fix was to invert the flow: before generating code, we generate context. We documented the implicit conventions — how we name services, where shared types live, which error patterns we use — in files the tools read every session. AI is only as good as the context you give it.</p>
<blockquote>
<p>💡 Team rule: if a pattern gets corrected twice in code review, it gets documented for the AI. The third time shouldn&#39;t exist.</p>
</blockquote>
<h2>What we automated (and what we didn&#39;t)</h2>
<p>Not all work benefits equally. Our split ended up like this:</p>
<ul>
<li><strong>We automate:</strong> module scaffolding, edge-case unit tests, repetitive migrations, mechanical cross-package refactors, and the first version of any CRUD.</li>
<li><strong>We assist:</strong> API design, business logic and complex queries — AI proposes, the human decides.</li>
<li><strong>We never delegate:</strong> architecture decisions, security, permissions and anything touching money or personal data.</li>
</ul>
<h2>Code review changed shape</h2>
<p>Reviewing AI-generated code is not like reviewing human code. AI code looks good — it&#39;s formatted, sensibly named and commented. The danger is in what looks correct. So we shifted the focus of review: less style, more behavior.</p>
<pre><code class="language-text">// Review checklist for AI-assisted code
// 1. Are the edge cases real or invented?
// 2. Does it reuse what exists or duplicate a helper?
// 3. Does error handling follow our pattern?
// 4. Are there tests that fail if behavior changes?
// 5. Did it touch anything outside the ticket&#39;s scope?
</code></pre>
<p>Point 5 turned out to be the most important: AI tools tend to &quot;improve&quot; neighboring code nobody asked them to touch. A clean, scoped diff is worth more than a brilliant, sprawling one.</p>
<h2>What I&#39;d do differently today</h2>
<p>I&#39;d start with context documentation from day one, not after the first incident. And I&#39;d set one simple metric that is, in the end, the only one that matters: not how much code the team generates, but how much code survives six months in production without being rewritten.</p>
<p>AI made us faster, yes. But the real change was something else: it forced us to make explicit all the knowledge that used to be tribal. And that would have been worth it even without the speed.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Painless serverless: Amplify, Lambda and DynamoDB for real apps</title>
      <link>https://yohangel.com/en/blog/serverless-sin-dolor/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/serverless-sin-dolor/</guid>
      <description>The architecture behind the Toyota Colombia admin panel: decisions, trade-offs, and what I would change today.</description>
      <pubDate>Thu, 28 May 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>AWS</category>
      <content:encoded><![CDATA[<p><strong>When we built the Toyota Colombia admin panel at Destiny, the question was never &quot;serverless or not?&quot; — it was &quot;how do we run a backend nobody has to babysit at 3 a.m.?&quot; The panel fed an offline-first PWA built in Next.js: catalogs, content, data that dealerships pulled from the field. Traffic was spiky, there was no ops team to speak of, and paying for servers running 24/7 to handle occasional bursts made no sense. Serverless fit because of the shape of the problem, not because it was trendy.</strong></p>
<h2>Why serverless made sense here</h2>
<p>I don&#39;t reach for serverless by default. I reach for it when the load pattern is spiky and the team is small. This project was both.</p>
<ul>
<li><strong>Bursty traffic.</strong> The panel was used in specific windows: content uploads, catalog updates, the odd query. In between, almost nothing. Paying for an EC2 box to sit idle for that is burning money.</li>
<li><strong>No platform team.</strong> There were no SREs. With Lambda and DynamoDB I don&#39;t patch operating systems, I don&#39;t manage scaling, and I don&#39;t get paged because a disk filled up.</li>
<li><strong>Per-event isolation.</strong> Every request is an invocation. A spike doesn&#39;t take the whole service down; it scales up and then back to zero.</li>
</ul>
<p>The price you pay is elsewhere: model DynamoDB well from day one, keep an eye on cold starts, and accept that you&#39;re married to AWS. It&#39;s worth it if you&#39;re honest about those costs.</p>
<h2>Modeling DynamoDB: single-table and access patterns</h2>
<p>The most common mistake with DynamoDB is treating it like Postgres without joins. It isn&#39;t. In DynamoDB you model <strong>the access patterns first</strong>, and the table second.</p>
<p>For the panel I used a single-table design: one table, generic <code>PK</code> and <code>SK</code>, several entity types living side by side. The patterns we needed were clear: fetch a vehicle by id, list vehicles by category, fetch the content tied to a model.</p>
<pre><code class="language-json">{
  &quot;PK&quot;: &quot;VEHICLE#corolla-2023&quot;,
  &quot;SK&quot;: &quot;METADATA&quot;,
  &quot;type&quot;: &quot;vehicle&quot;,
  &quot;name&quot;: &quot;Corolla&quot;,
  &quot;year&quot;: 2023,
  &quot;category&quot;: &quot;sedan&quot;,
  &quot;GSI1PK&quot;: &quot;CATEGORY#sedan&quot;,
  &quot;GSI1SK&quot;: &quot;VEHICLE#corolla-2023&quot;
}
</code></pre>
<p>The trick: fetch-by-id goes against <code>PK</code>/<code>SK</code>, and list-by-category goes against a global secondary index (<code>GSI1PK</code>/<code>GSI1SK</code>). No full-table scans. Every read is a <code>Query</code> against a known partition key, which is the only thing DynamoDB does cheaply and predictably.</p>
<p>Rules that saved me pain:</p>
<ol>
<li><strong>Write down the list of access patterns before touching the table.</strong> If a new one shows up later, it&#39;s almost always one more GSI, not a redesign.</li>
<li><strong>Prefix your keys</strong> (<code>VEHICLE#</code>, <code>CATEGORY#</code>) so different entities coexist without colliding.</li>
<li><strong>Never a <code>Scan</code> on the hot path.</strong> If you need to filter by something, that something belongs in a key.</li>
</ol>
<h2>Cold starts: real, but tameable</h2>
<p>Cold starts exist and they&#39;re annoying, but in 2026 the drama is overblown. In this panel, the backend was Node functions on Lambda behind API Gateway, and the extra latency on the first invocation was never a business problem: this was an internal tool, not a checkout with a millisecond SLA.</p>
<p>That said, here&#39;s what I actually do:</p>
<ul>
<li><strong>Keep the bundle small.</strong> Fewer dependencies to load means a faster start. Bundling with esbuild and dropping what I don&#39;t use matters more than any clever trick.</li>
<li><strong>Instantiate the DynamoDB client outside the handler,</strong> so the connection is reused across warm invocations.</li>
<li><strong>Memory as a CPU lever.</strong> On Lambda, CPU scales with allocated memory; bumping from 128 to 512 MB is often cheaper <em>and</em> faster, because the function finishes sooner.</li>
</ul>
<pre><code class="language-ts">import { DynamoDBClient } from &quot;@aws-sdk/client-dynamodb&quot;;
import { DynamoDBDocumentClient, QueryCommand } from &quot;@aws-sdk/lib-dynamodb&quot;;

// Outside the handler: reused across warm invocations.
const ddb = DynamoDBDocumentClient.from(new DynamoDBClient({}));

export const handler = async (event: { category: string }) =&gt; {
  const res = await ddb.send(
    new QueryCommand({
      TableName: process.env.TABLE_NAME,
      IndexName: &quot;GSI1&quot;,
      KeyConditionExpression: &quot;GSI1PK = :pk&quot;,
      ExpressionAttributeValues: { &quot;:pk&quot;: `CATEGORY#${event.category}` },
    }),
  );
  return { items: res.Items ?? [] };
};
</code></pre>
<blockquote>
<p>💡 With serverless you don&#39;t pay for servers — you pay for designing your access patterns right on day one. That design debt doesn&#39;t refactor cheaply.</p>
</blockquote>
<h2>Amplify for the frontend and auth</h2>
<p>The panel itself shipped on Amplify: frontend hosting, auth with Cognito underneath, and the glue between the client and the Lambdas. For an internal panel with a login, Amplify saves you weeks of plumbing. You don&#39;t stand up the session flow or the CI-backed hosting yourself.</p>
<p>It&#39;s convenient, at a price: Amplify is opinionated. The moment you step off its happy path, the abstraction starts to weigh on you, and debugging what it generates underneath gets tedious. For the scope of this project, the trade was a good one.</p>
<h2>Infrastructure as code and the ops trade-offs</h2>
<p>All of this lives in code. No hand-clicking resources in the console: tables, functions, IAM roles — all declared. Today, with the experience of running JXBS&#39;s infra in OpenTofu (ECS Fargate, queues, Secrets Manager, CI/CD), I&#39;m even more convinced that <strong>the worst mistake is having resources nobody knows the origin of.</strong> If it isn&#39;t in code, it doesn&#39;t exist.</p>
<p>The upside of this model, no dressing it up:</p>
<ul>
<li>No servers to patch or maintain.</li>
<li>Scale to zero: no traffic, no compute bill.</li>
<li>A smaller ops surface for a small team.</li>
</ul>
<p>The uncomfortable part, with the same honesty:</p>
<ul>
<li><strong>Real lock-in.</strong> DynamoDB and Lambda don&#39;t migrate to another cloud without a rewrite.</li>
<li><strong>Debugging is different.</strong> There&#39;s no server to SSH into; you live in logs and traces.</li>
<li><strong>Cost can surprise you</strong> if you <code>Scan</code> or model badly: DynamoDB is dirt cheap used well and brutal used badly.</li>
</ul>
<h2>What I&#39;d do differently today (2026)</h2>
<p>I&#39;d pick serverless again for this case; the shape of the problem hasn&#39;t changed. But I&#39;d change a few things:</p>
<ul>
<li><strong>Less Amplify, more explicit control.</strong> Today I&#39;d wire up hosting and auth more directly and keep Amplify for prototypes, not for something that has to live for years.</li>
<li><strong>Start in OpenTofu, not the proprietary tool.</strong> It&#39;s what I run in production now, and I wouldn&#39;t go back.</li>
<li><strong>Invest in observability earlier.</strong> Distributed tracing from day one; in serverless, not having it is flying blind.</li>
</ul>
<p>Painless serverless doesn&#39;t mean serverless without decisions. It means making the hard decisions up front — data modeling, IaC, observability — so you&#39;re not paying for them at 3 a.m. later.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Offline-first PWAs: apps that work when the signal does not</title>
      <link>https://yohangel.com/en/blog/pwas-offline-first/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/pwas-offline-first/</guid>
      <description>Service Workers, smart caching and deferred sync for events with spotty connectivity.</description>
      <pubDate>Fri, 15 May 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>Frontend</category>
      <content:encoded><![CDATA[<p><strong>At a Toyota event in Colombia, connectivity isn&#39;t a guaranteed luxury: it&#39;s a packed hall, saturated mobile networks, and booths where the WiFi comes and goes. That&#39;s where I built, with the Destiny team, a Next.js PWA that didn&#39;t depend on the signal to work. Not a &quot;network-fault-tolerant&quot; app, but one designed from day one to operate offline and reconcile once the signal came back. That difference in mindset —offline-first, not offline-tolerant— changed how we designed storage, sync, and above all how we surfaced state to non-technical people who were busy attending to customers right there on the floor.</strong></p>
<h2>Offline-first is not the same as offline-tolerant</h2>
<p>An offline-tolerant app assumes the network is the norm and offline is the exception: when it fails, it shows a friendly error and waits. An offline-first app flips that premise. The immediate source of truth is local; the network is an implementation detail used to propagate and refresh data when it happens to be available.</p>
<p>At the Toyota event this wasn&#39;t an aesthetic choice. A rep capturing an interested customer&#39;s details can&#39;t sit staring at a spinner because the hall has saturated the mobile network. The capture had to complete every time, save locally, and go out to the server whenever it could. The backend was serverless on AWS —Amplify, Lambda, and DynamoDB— but from the frontend&#39;s point of view, the server might not exist for minutes and the app stayed fully usable.</p>
<h2>Caching strategies in the service worker</h2>
<p>The service worker is the heart of an offline-first PWA. The key is not to use a single strategy, but to pick one per resource type:</p>
<ul>
<li><strong>App shell (precache):</strong> HTML, JS, CSS, and shell assets are precached on install. The app boots with no network.</li>
<li><strong>Stale-while-revalidate:</strong> for resources that can be slightly out of date (catalog images, icons). I serve the cache instantly and refresh in the background.</li>
<li><strong>Network-first with cache fallback:</strong> for data I want fresh when there&#39;s signal, but that can&#39;t block the experience when there isn&#39;t.</li>
</ul>
<pre><code class="language-js">// sw.js — routed by resource type
self.addEventListener(&#39;fetch&#39;, (event) =&gt; {
  const { request } = event;
  const url = new URL(request.url);

  // API data: network-first with cache fallback
  if (url.pathname.startsWith(&#39;/api/&#39;)) {
    event.respondWith(
      fetch(request)
        .then((res) =&gt; {
          const copy = res.clone();
          caches.open(&#39;api-cache&#39;).then((c) =&gt; c.put(request, copy));
          return res;
        })
        .catch(() =&gt; caches.match(request))
    );
    return;
  }

  // Catalog assets: stale-while-revalidate
  if (url.pathname.startsWith(&#39;/catalog/&#39;)) {
    event.respondWith(
      caches.match(request).then((cached) =&gt; {
        const network = fetch(request).then((res) =&gt; {
          caches.open(&#39;catalog-cache&#39;).then((c) =&gt; c.put(request, res.clone()));
          return res;
        });
        return cached || network;
      })
    );
    return;
  }

  // App shell: cache-first
  event.respondWith(caches.match(request).then((c) =&gt; c || fetch(request)));
});
</code></pre>
<h2>Pending writes in IndexedDB</h2>
<p>Reading offline is easy; writing offline is where the real work lives. Every data capture a rep made was first saved to IndexedDB as a pending operation, with its own client-generated <code>id</code> (a UUID), a timestamp, and the full payload. The UI gave immediate feedback: the record was saved, full stop. It didn&#39;t wait on the server.</p>
<p>That local queue is what turns offline into something you can trust. Nothing lives only in memory; if the rep closed the tab or ran out of battery, the operation was still there on reopen.</p>
<blockquote>
<p>💡 In offline-first the network isn&#39;t the source of truth, it&#39;s just a sync channel. If your UI waits on the server to confirm anything, you&#39;re still building an offline-tolerant app in disguise.</p>
</blockquote>
<h2>Deferred sync when the signal returns</h2>
<p>When connectivity came back, that queue had to be drained. I used the Background Sync API to hand that work to the browser: I register a sync event and the system itself decides when there&#39;s a stable network to run it, even if the tab is no longer in the foreground.</p>
<pre><code class="language-js">// When queuing a write, ask for a deferred sync
async function queueWrite(record) {
  await idbPut(&#39;pending-writes&#39;, record); // save to IndexedDB
  const reg = await navigator.serviceWorker.ready;
  if (&#39;sync&#39; in reg) {
    await reg.sync.register(&#39;flush-writes&#39;);
  }
}

// In the service worker: drain the queue when the browser allows it
self.addEventListener(&#39;sync&#39;, (event) =&gt; {
  if (event.tag === &#39;flush-writes&#39;) {
    event.waitUntil(flushPendingWrites());
  }
});

async function flushPendingWrites() {
  const pending = await idbGetAll(&#39;pending-writes&#39;);
  for (const record of pending) {
    const res = await fetch(&#39;/api/leads&#39;, {
      method: &#39;POST&#39;,
      headers: {
        &#39;Content-Type&#39;: &#39;application/json&#39;,
        &#39;Idempotency-Key&#39;: record.id, // the client UUID
      },
      body: JSON.stringify(record),
    });
    if (res.ok) await idbDelete(&#39;pending-writes&#39;, record.id);
  }
}
</code></pre>
<h2>Conflicts and idempotency</h2>
<p>The scenario you can&#39;t ignore: the write did reach the server, but the response got lost on the network before coming back. The client thinks it failed and retries. Without protection, you duplicate the record.</p>
<p>The fix was idempotency using the client&#39;s key. That UUID generated when the data was captured traveled as the <code>Idempotency-Key</code>. In the Lambda, if a record with that key already existed, we returned the existing result instead of creating a new one. Retrying was safe by design: a thousand retries produced a single record.</p>
<p>For edit conflicts on the same piece of data, the rule was simple and explainable: last-write-wins by client timestamp. It wasn&#39;t CRDTs or anything sophisticated, but for the event&#39;s flow —mostly independent captures— it was enough and predictable.</p>
<h2>The UX of sync state</h2>
<p>Everything above is invisible to the rep, and that&#39;s how it should be almost always. But &quot;almost&quot; is the operative word. The person at the booth needs confidence, not technical detail. We showed three simple states:</p>
<ul>
<li><strong>Saved</strong> (local, not yet synced): a check and &quot;saved on device.&quot;</li>
<li><strong>Syncing:</strong> a discreet indicator while the queue was draining.</li>
<li><strong>Synced:</strong> confirmation the data now lived on the server.</li>
</ul>
<p>I also showed a &quot;pending to send&quot; counter. That reassured a non-technical person: they knew nothing had been lost even without signal, and they knew when everything was safe before closing out their shift.</p>
<h2>What I&#39;d do differently in 2026</h2>
<p>Today I&#39;d rethink a couple of things. First, I&#39;d lean more on mature libraries instead of writing so much by hand: Workbox for the service worker strategies, and a layer like Dexie over IndexedDB instead of the raw API. Second, for sync I&#39;d evaluate a real local-first engine —something with CRDT-based reconciliation when the flow justifies it— so I&#39;m not locked into last-write-wins in cases with genuine concurrent editing. And third, I&#39;d be more deliberate about observability: record queue metrics (size, operation age, retry rate) to diagnose in the field, because at an event you don&#39;t have time to open DevTools. The core idea, though, I wouldn&#39;t touch: the signal is optional, the work is not.</p>
]]></content:encoded>
    </item>
    <item>
      <title>Embeddings + pgvector: semantic matching in PostgreSQL</title>
      <link>https://yohangel.com/en/blog/embeddings-pgvector/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/embeddings-pgvector/</guid>
      <description>How we built the JXBS matching engine without adding a dedicated vector database.</description>
      <pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>IA</category>
      <content:encoded><![CDATA[<p><strong>When I built the matching engine at JXBS, the first decision wasn&#39;t which embedding model to use — it was where to store the vectors. And I chose not to add a dedicated vector database. The embeddings live in the same PostgreSQL that already holds everything else: candidates, openings, applications. With the pgvector extension, Postgres stores, indexes, and searches by similarity without me having to keep two systems in sync. This article is the reasoning behind that call and how I actually built it.</strong></p>
<h2>Why I stayed in Postgres</h2>
<p>The pull toward Pinecone, Weaviate, or Qdrant is real. But every new datastore is one more thing to operate, monitor, back up, and keep consistent. In a recruitment product, semantic similarity never travels alone: I always cross it with structured data. Location, seniority, availability, whether the opening is still active. If the vectors live in another system, every search becomes two queries and a join stitched together by hand in application code.</p>
<p>Keeping it all in Postgres buys me three concrete things. Transactions: when I update a candidate and their embedding, either both land or neither does. Joins: I filter on structured columns and order by vector distance in the same query. And operational simplicity: one backup, one place to monitor, one Prisma migration. On ECS Fargate with OpenTofu, every extra piece of infrastructure is paid for in maintenance time.</p>
<h2>How embeddings work, at a practical level</h2>
<p>An embedding is a vector of numbers representing the meaning of a piece of text. Similar texts land close together in that space; unrelated ones land far apart. I don&#39;t need to understand high-dimensional geometry to use it: I hand the text to an embedding model, it returns an array of floats, and I store it.</p>
<p>What matters is which text I feed it. For a candidate I don&#39;t pass the raw résumé wholesale — I build a structured summary of experience, technologies, and role. For an opening, the title plus the actual requirements. Garbage in, garbage out: embedding quality depends on the quality of the source text far more than on the model.</p>
<h2>Storing vectors with pgvector</h2>
<p>pgvector adds a <code>vector</code> type and distance operators. The table looks like this:</p>
<pre><code class="language-sql">CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE candidate_embedding (
  candidate_id  UUID PRIMARY KEY REFERENCES candidate(id),
  embedding     vector(1536) NOT NULL,
  seniority     TEXT NOT NULL,
  location      TEXT NOT NULL,
  is_available  BOOLEAN NOT NULL DEFAULT true,
  updated_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE INDEX ON candidate_embedding
  USING hnsw (embedding vector_cosine_ops)
  WITH (m = 16, ef_construction = 64);
</code></pre>
<p>The <code>&lt;=&gt;</code> operator is cosine distance. Smaller distance, higher similarity. I pick cosine because I care about the vector&#39;s direction, not its magnitude.</p>
<h2>Index choice: HNSW vs IVFFlat</h2>
<p>pgvector offers two index types and the difference matters. IVFFlat groups vectors into lists and searches only the nearest ones: it builds fast and stays small, but its recall depends on how many lists you probe and it struggles as the data grows. HNSW builds a layered navigable graph: better recall at lower query latency, in exchange for slower builds and higher memory use.</p>
<p>I went with HNSW. In recruitment, a relevant candidate who never surfaces is a silent failure — nobody reports it, but it quietly degrades the product. I&#39;d rather pay in index build time and RAM for stable recall. <code>ef_construction</code> controls graph quality at build time; at query time, raising <code>ef_search</code> improves recall at the cost of latency. That&#39;s the lever for the trade-off, and I move it based on what I measure, not on a hunch.</p>
<blockquote>
<p>💡 Cosine distance tells you what&#39;s similar; it doesn&#39;t tell you what&#39;s right. That gap is the whole product.</p>
</blockquote>
<h2>Hybrid search: why I don&#39;t trust cosine alone</h2>
<p>This is the point that took me longest to learn. Semantic similarity is great at finding things that resemble each other, but &quot;resembles&quot; isn&#39;t &quot;correct.&quot; A senior backend engineer in Caracas and a junior one in another time zone can have close embeddings because they share technologies. Cosine has no idea the opening demands high seniority and local presence.</p>
<p>So I run hybrid search: vector similarity is a signal, not the verdict. I filter with <code>WHERE</code> on structured columns and blend the distance with a structured score.</p>
<pre><code class="language-sql">SELECT
  c.candidate_id,
  1 - (c.embedding &lt;=&gt; $1::vector)              AS similarity,
  CASE WHEN c.seniority = $2 THEN 0.3 ELSE 0 END AS seniority_boost
FROM candidate_embedding c
WHERE c.is_available = true
  AND c.location = $3
ORDER BY (c.embedding &lt;=&gt; $1::vector)
         - CASE WHEN c.seniority = $2 THEN 0.3 ELSE 0 END
LIMIT 20;
</code></pre>
<p>The <code>WHERE</code> trims the space to eligible candidates before ranking. Semantic similarity orders within that set, and a structured boost adjusts the final order. The weights are configurable, and I tune them against what the recruitment team considers a good match, not against an abstract cosine metric.</p>
<h2>Chunking and keeping embeddings fresh</h2>
<p>Long résumés don&#39;t go in as a single pass. I split them into meaningful sections (experience, skills, education) so the &quot;backend experience&quot; embedding isn&#39;t diluted by three paragraphs of hobbies. It&#39;s pragmatic chunking, guided by document structure rather than a fixed token size.</p>
<p>Freshness: an embedding is a snapshot of the text at a moment in time. If a candidate updates their profile or an opening changes its requirements, the vector goes stale. I store a hash of the source text; when it changes, I re-enqueue regeneration. That way I only recompute what actually changed, and I don&#39;t burn model calls for nothing.</p>
<h2>Costs, and when I would reach for a vector DB</h2>
<p>Generating embeddings costs per token, and that&#39;s the real spend — not storage. Re-enqueuing only on a hash change keeps the bill in check. HNSW is heavy in RAM, so I size the Postgres instance with the index in mind, not just the rows.</p>
<p>When would I leave Postgres? If I hit hundreds of millions of vectors with very high, sustained search traffic, where the index no longer fits comfortably in memory and latency becomes the business bottleneck. At that scale, a dedicated vector engine with sharding earns its keep. But that&#39;s a decision you make with real load numbers, not by front-running a problem that may never arrive. Until then, Postgres with pgvector does the job and leaves me with fewer moving parts that can break.</p>
]]></content:encoded>
    </item>
    <item>
      <title>From monolith to per-audience backends: migration lessons</title>
      <link>https://yohangel.com/en/blog/monolito-a-backends/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/monolito-a-backends/</guid>
      <description>Why we split a shared NestJS backend, and the pnpm DI pitfalls nobody warns you about.</description>
      <pubDate>Sat, 18 Apr 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>AWS</category>
      <content:encoded><![CDATA[<p><strong>For almost three years, a single NestJS backend served us well. One monolith, one deploy, one mental model. Until it didn&#39;t. There was no dramatic outage, no production incident: it was the daily friction of touching one portal&#39;s code and having to think about the other three. At JXBS we build an AI-first recruitment SaaS with several portals for distinct audiences —candidates, companies, internal admin— and they all lived inside the same backend. This is the story of why we split it into per-audience backends, and the things nobody warns you about regarding pnpm and NestJS dependency injection when you do.</strong></p>
<h2>Why one backend stopped scaling with us</h2>
<p>The problem was never performance. It was coupling and the blast radius of every deploy.</p>
<p>When everything lives under one NestJS root module, the boundaries blur on their own. A candidates service starts importing something from the companies module &quot;just for now.&quot; A guard meant for admin sneaks into a public route. And since it&#39;s a single process, everything shares the same config, the same PostgreSQL connection pool, the same startup cycle.</p>
<p>Three concrete symptoms pushed us to move:</p>
<ul>
<li><strong>Deploy blast radius.</strong> A change in the admin portal forced a redeploy of the entire backend. If something broke, it all went down, including the public candidates portal that hadn&#39;t changed in weeks.</li>
<li><strong>Mixed audiences.</strong> Different portals have different auth requirements, different rate limits, different exposure surface. Cramming them into one process meant the most sensitive route set the paranoia level for everything.</li>
<li><strong>Shared cognition.</strong> Every person on the team had to load the whole domain into their head just to touch a small part of it.</li>
</ul>
<h2>What &quot;per-audience backends&quot; means</h2>
<p>The idea is simple: one backend per portal. One for candidates, one for companies, one for admin. Each is its own independent NestJS 11 service, with its own deploy, its own config, and its own API surface. The frontend stays Next.js 15 with React 19, and each portal talks to its backend.</p>
<p>What does <strong>not</strong> change is the database or the domain. We&#39;re still on PostgreSQL + pgvector with Prisma. The trick is that shared code —entities, domain logic, utilities, the Prisma client— lives in monorepo packages (Turborepo + pnpm), not duplicated across services.</p>
<pre><code class="language-text">apps/
  candidates-api/        # NestJS 11 — candidates portal
  companies-api/         # NestJS 11 — companies portal
  admin-api/             # NestJS 11 — admin portal
  web-candidates/        # Next.js 15
  web-companies/         # Next.js 15
packages/
  domain/                # pure domain logic, no NestJS
  database/              # Prisma client + repos
  auth/                  # reusable auth module
  config/                # env loading + validation
</code></pre>
<p>The rule that saved us: packages under <code>packages/</code> expose classes and functions; the <code>apps/</code> decide how to wire them. The domain doesn&#39;t know which backend it&#39;s running in.</p>
<h2>The pnpm + dependency-injection traps nobody mentions</h2>
<p>This is where I lost hours. NestJS assumes a provider is a singleton within its container. pnpm, with its strict <code>node_modules</code> layout and its symlinks, can quietly break that assumption.</p>
<p>The first hit: <strong>duplicated provider instances</strong>. If a shared package declares a provider and two different modules import it via slightly different paths, you end up with two instances. A &quot;singleton&quot; that isn&#39;t one. You notice it when an in-memory cache or a connection pool behaves inconsistently.</p>
<p>The cause is almost always the same: the provider token isn&#39;t identical across importers, or the shared dependency got installed in two places in the tree.</p>
<p>The pattern that avoids it: a <strong>global dynamic module</strong> with an explicit token, declared exactly once.</p>
<pre><code class="language-ts">// packages/database/src/database.module.ts
import { Global, Module, DynamicModule } from &#39;@nestjs/common&#39;;
import { PrismaService } from &#39;./prisma.service&#39;;

export const PRISMA = Symbol(&#39;PRISMA&#39;); // stable, unique token

@Global()
@Module({})
export class DatabaseModule {
  static forRoot(): DynamicModule {
    return {
      module: DatabaseModule,
      providers: [{ provide: PRISMA, useClass: PrismaService }],
      exports: [PRISMA],
    };
  }
}
</code></pre>
<p>Each app calls <code>DatabaseModule.forRoot()</code> <strong>once</strong> in its root module. The token is an exported <code>Symbol</code>, so there&#39;s no string ambiguity and no collision risk.</p>
<p>The other two that bit me:</p>
<ul>
<li><strong>Peer deps, not direct deps.</strong> Shared packages declare <code>@nestjs/common</code> and <code>@nestjs/core</code> as <code>peerDependencies</code>, never as regular dependencies. If a package ships its own copy of NestJS, its decorators and metadata live in a different realm and DI simply won&#39;t recognize them. In <code>pnpm-workspace.yaml</code>, leaning on the catalog to pin a single version helps a lot.</li>
<li><strong>Module boundaries, not folder boundaries.</strong> A package can export code; whether it&#39;s a NestJS <code>@Module</code> is a separate decision. We kept <code>packages/domain</code> NestJS-free —pure classes and functions— and only <code>packages/auth</code> and <code>packages/database</code> expose modules. That way the domain gets reused without dragging the DI container everywhere.</li>
</ul>
<blockquote>
<p>💡 In a monorepo, a &quot;singleton&quot; is only a singleton if the token is identical and the dependency is installed exactly once. pnpm won&#39;t warn you when that&#39;s false: it tells you in production.</p>
</blockquote>
<h2>How to share domain logic without recoupling it</h2>
<p>The obvious risk of splitting the monolith is rebuilding the coupling inside the packages. A <code>shared</code> package that knows everything is the monolith with a new name.</p>
<p>What worked:</p>
<ul>
<li><strong>The domain imports no infrastructure.</strong> <code>packages/domain</code> knows nothing about Prisma or NestJS. It takes interfaces; the apps inject implementations.</li>
<li><strong>One package, one responsibility.</strong> <code>auth</code>, <code>database</code>, <code>config</code> kept separate. If you&#39;re unsure which package something belongs to, it probably belongs to the app.</li>
<li><strong>No cross-backend dependencies.</strong> <code>candidates-api</code> never imports from <code>companies-api</code>. If they need to talk, it&#39;s over an API or a shared contract in <code>packages/</code>, never a direct import.</li>
</ul>
<h2>Deploying multiple services on ECS Fargate</h2>
<p>Each backend is its own service on ECS Fargate: its own task definition, its own image, its own auto-scaling. We manage CI/CD with OpenTofu, so adding a backend is declaring a new module, not clicking through a console.</p>
<p>Turborepo gives us the other pillar: affected builds. Only what changed gets built and deployed. Touching <code>admin-api</code> doesn&#39;t rebuild <code>candidates-api</code>. That was the original goal —shrinking the blast radius— and this is where it becomes real.</p>
<p>The detail that watches the budget: each service sizes its CPU and memory to its own load. The candidates portal and the admin portal don&#39;t have to pay for the same task size.</p>
<h2>Doing it without a big-bang</h2>
<p>We rewrote nothing all at once. We extracted one backend first —the lowest-risk one— leaving the monolith serving the rest. Portal by portal, we moved routes to the new service and pointed the corresponding frontend at it. The monolith kept shrinking until it was gone.</p>
<h2>What I&#39;d tell my past self</h2>
<p>Draw the audience boundaries from day one, even if everything lives in a single process at first. Separating modules with clear limits is cheap; untangling a coupled monolith is expensive. And treat dependency injection in a pnpm monorepo as what it really is: a question of instance identity, not of imports. Pin the versions, use explicit tokens, keep the domain pure. The rest follows.</p>
]]></content:encoded>
    </item>
    <item>
      <title>The design system as a product: maintaining a shared UI library</title>
      <link>https://yohangel.com/en/blog/design-system-producto/</link>
      <guid isPermaLink="true">https://yohangel.com/en/blog/design-system-producto/</guid>
      <description>Versioning, tokens and component discipline when several portals depend on your library.</description>
      <pubDate>Thu, 02 Apr 2026 00:00:00 GMT</pubDate>
      <dc:creator>Yohangel Ramos</dc:creator>
      <category>Frontend</category>
      <content:encoded><![CDATA[<p><strong>For a long time I treated our component library like a shared folder: a place to drop buttons and modals so we wouldn&#39;t rewrite them. The day several portals in the monorepo started depending on it at the same time, I realized it wasn&#39;t a folder anymore. It was a product, with real users —the rest of the team— who feel every careless change. Since then I&#39;ve maintained it as Tech Lead at JXBS with that mindset: my users are other developers, and their experience matters as much as the end user&#39;s.</strong></p>
<h2>Tokens are the contract, not a detail</h2>
<p>Before you talk about components you have to talk about tokens, because they&#39;re the first thing you break without noticing. A token is a design decision with a stable name: a color, a spacing step, a type scale. If one portal hardcodes <code>#2563eb</code> and another uses <code>var(--color-primary)</code>, you don&#39;t have a design system, you have two systems that happen to match by accident.</p>
<p>I treat tokens as the public contract. Consumers don&#39;t ask for hex colors; they ask for semantic intent.</p>
<pre><code class="language-css">:root {
  /* Primitives: never used directly in components */
  --blue-600: #2563eb;
  --gray-900: #111827;
  --space-2: 0.5rem;
  --space-4: 1rem;

  /* Semantic: this is what the consumer touches */
  --color-action: var(--blue-600);
  --color-text-strong: var(--gray-900);
  --radius-control: 0.5rem;
}

[data-theme=&quot;dark&quot;] {
  --color-text-strong: #f9fafb;
}
</code></pre>
<p>The semantic layer is what lets me theme without touching a single component. I change the mapping, not the signature. And that split between primitives and semantics is exactly what stops someone from coupling their portal to a specific blue that I&#39;ll want to move tomorrow.</p>
<h2>A component&#39;s API is a promise</h2>
<p>With React 19 and TypeScript, a component&#39;s signature is a contract as serious as any HTTP API. Every prop I expose is something I&#39;ll have to hold up across versions. That&#39;s why I lean on composition over configuration: instead of a <code>&lt;Card&gt;</code> with twenty boolean props, I expose pieces that combine.</p>
<pre><code class="language-tsx">// Bad: configuration that grows without end
&lt;Card title=&quot;...&quot; footer withBorder elevated dense hasIcon /&gt;

// Better: composition, the consumer assembles their case
&lt;Card&gt;
  &lt;Card.Header&gt;...&lt;/Card.Header&gt;
  &lt;Card.Body&gt;...&lt;/Card.Body&gt;
  &lt;Card.Footer&gt;...&lt;/Card.Footer&gt;
&lt;/Card&gt;
</code></pre>
<p>Two rules I don&#39;t negotiate. First: don&#39;t leak the internal DOM. If I expose a <code>className</code> that lands on some arbitrary <code>&lt;div&gt;</code>, and tomorrow that <code>&lt;div&gt;</code> disappears, I break everyone who relied on that structure. I&#39;d rather offer explicit extension points (<code>asChild</code>, slots, typed variants) than let people style my guts. Second: stable props are sacred. Renaming <code>variant</code> to <code>kind</code> &quot;because it reads better&quot; isn&#39;t a refactor, it&#39;s breaking six teams at once.</p>
<blockquote>
<p>💡 A design system isn&#39;t measured by how many components it has, but by how many people can build on top of it without messaging you.</p>
</blockquote>
<h2>Versioning inside the monorepo without hurting people</h2>
<p>Turborepo and pnpm let everything live together, and that has a trap: it&#39;s far too easy to &quot;bump everything&quot; at once. I use changesets with real semver. A semantic color change is a patch. A new optional prop is a minor. Removing a prop, renaming it, or changing DOM that someone might be selecting is a major —and a major demands notice, a migration guide, and sometimes a window where the old version keeps living alongside the new one.</p>
<p>What I religiously avoid is the reflex &quot;bump everything.&quot; If I only touch <code>Button</code>, I don&#39;t want to force a portal that doesn&#39;t even use it to rebuild and re-verify. Granular versioning is what keeps trust alive: when you upgrade, you know what&#39;s going to hurt before it hurts.</p>
<h2>Accessibility for free, and the social side</h2>
<p>At Acid Tango, in Madrid, I pushed WCAG 2.1 AA as a team standard, and that experience changed how I build libraries. The best accessibility is the kind the consumer gets for free. If my <code>Dialog</code> handles focus, <code>Escape</code>, <code>aria-modal</code>, and returning focus to the trigger, then every portal is accessible without its developer knowing anything about ARIA. Keyboard support, focus management, and contrast aren&#39;t optional features: they&#39;re the floor.</p>
<p>But the hardest part isn&#39;t technical, it&#39;s social. A design system has governance whether anyone writes it down or not. I try to make it explicit: who decides what gets in, how you propose a new component, and —the uncomfortable one— how you say no. I say no when something serves a single portal; that lives in the portal, not in the core. To onboard a new component I ask for three things: a real, repeated use case, documentation with copy-pasteable examples, and accessibility solved before it merges. Without examples, people reinvent; and every reinvention is a crack in the contract.</p>
<h2>What I&#39;d do differently in 2026</h2>
<p>I&#39;d document before coding. For a while documentation showed up afterward, and that&#39;s too late: if you can&#39;t explain the API in a short example, the API is wrong. Today I&#39;d write the usage example first and let it dictate the signature.</p>
<p>I&#39;d also be more aggressive about graduated deprecation. Marking a prop <code>@deprecated</code> in TypeScript, letting it keep working for a few versions, and warning on every build is infinitely kinder than a surprise major. A mature design system isn&#39;t measured by how fast it changes, but by how predictable it is when it does.</p>
]]></content:encoded>
    </item>
  </channel>
</rss>
