<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>I Made My Local AI Pick Its Own Model</title>
        <link>https://tube.blueben.net/videos/watch/cbf10e6c-cb58-48f3-a33d-2cf7b428429f</link>
        <description>I wanted my local AI to choose the right model automatically instead of making me pick between speed and capability for every request. I test NVIDIA Switchyard for local AI model routing with Nemotron 3.5 Lightning on an RTX 3090 and DeepSeek V4 Flash across two GX10s. I cover Stage routing, vLLM, context windows, and some of the problems I ran into while testing it with real coding agents. In my best autonomous coding run, almost 88% of the routing decisions went to Nemotron while DeepSeek was still available when Switchyard decided it was needed. This is pretty close to what I wanted. Most of the work stays on the fast local model, the more capable model gets pulled in when it makes sense, and I don't have to choose between them before I start. Video Notes: https://technotim.com/posts/switchyard-local-ai-routing/ Merch Shop 🛍️: https://l.technotim.com/shop Support me on Patreon: https://www.patreon.com/technotim Sponsor me on GitHub: https://github.com/sponsors/timothystewart6 Subscribe on Twitch: https://www.twitch.tv/technotim Become a YouTube member: https://www.youtube.com/channel/UCOk-gHyjcWZNj3Br4oxwh0A/join Gear Recommendations: https://l.technotim.com/gear Get Help in Our Discord Community: https://l.technotim.com/discord 2nd channel: https://www.youtube.com/@TechnoTimTinkers (Affiliate links may be included in this description. I may receive a small commission at no cost to you.) 00:00 Automatic AI Model Routing 01:09 Nemotron Lightning vs DeepSeek 01:40 Fixing Nemotron's Performance 02:35 Is Nemotron Actually Good? 02:58 How NVIDIA Switchyard Works 04:55 Putting Stage Routing to the Test 05:28 The 32K Context Problem 06:46 Fixing Routing with 128K Context 07:44 Testing on Real-World Code 08:40 Building the Autonomous Agent Test 09:46 Run 4: Getting Close 10:39 Run 5: 88% Routed to Nemotron 11:50 Run 6: When the Agent Got Stuck 12:11 Is NVIDIA Switchyard Worth It? 13:40 What's Next for My Local AI Thank you for watching!</description>
        <lastBuildDate>Tue, 06 Oct 2026 17:25:30 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>PeerTube - https://tube.blueben.net</generator>
        <image>
            <title>I Made My Local AI Pick Its Own Model</title>
            <url>https://tube.blueben.net/client/assets/images/icons/icon-96x96.png</url>
            <link>https://tube.blueben.net/videos/watch/cbf10e6c-cb58-48f3-a33d-2cf7b428429f</link>
        </image>
        <copyright>All rights reserved, unless otherwise specified in the terms specified at https://tube.blueben.net/about and potential licenses granted by each content's rightholder.</copyright>
        <atom:link href="https://tube.blueben.net/feeds/video-comments.xml?videoId=cbf10e6c-cb58-48f3-a33d-2cf7b428429f" rel="self" type="application/rss+xml"/>
    </channel>
</rss>