← Retour au fil
Instant (SN46) : l'approche logicielle pour l'inférence analogique
TAO Daily13 août, 11h · il y a 1j

Instant (SN46) : l'approche logicielle pour l'inférence analogique

Sur Bittensor, le sous-réseau Instant (SN46) bouscule les standards : en exécutant son propre code sur le matériel des mineurs, il atteint des vitesses d'inférence IA bien au-delà d'OpenRouter.

Le sous-réseau Instant (SN46) sur Bittensor résout le problème de confiance de l'inférence IA sans recourir aux TEE. Via le système "Managed VNet IR", le code du protocole s'exécute directement sur le matériel des mineurs. Né d'un pivot depuis le projet Zipcode, Instant affiche des performances de 300 à 500 jetons par seconde sur GPT-OSS 20B, soit le double d'OpenRouter.

Cette architecture permet d'exploiter n'importe quel type de matériel, des GPU aux FPGA. Instant a déjà signé avec Ditto (SN118) comme premier client. L'objectif final est de graver les poids des modèles d'IA directement dans des puces en silicium analogique, une technologie poussée par des acteurs comme Taalas (racheté par AMD) pour des vitesses de traitement extrêmes.

BittensorSky

Détails

Source
TAO Daily
Publication
13 août à 11h54

Contenu source (brut)

<div id="bsf_rt_marker"></div> <p class="wp-block-paragraph">Most inference subnets on Bittensor solve the trust problem by wrapping miners in TEEs and verifying attestations. Instant (SN46) solves it by not needing miners to be trusted at all.</p> <figure class="wp-block-embed is-type-rich is-provider-x wp-block-embed-x"><div class="wp-block-embed__wrapper"> <div class="embed-x"><blockquote class="twitter-tweet" data-width="500" data-dnt="true"><p lang="en" dir="ltr">Hash Rate &#8211; Ep. 182: &#39;Instant&#39; Inference Subnet 46<br>🧙 Guest: <a href="https://x.com/opendansor?ref_src=twsrc%5Etfw">@opendansor</a> of <a href="https://x.com/instantsubnet?ref_src=twsrc%5Etfw">@instantsubnet</a> <br><br>00:00 Introduction to Dan and his background<br>07:50 ZK proofs and their application in AI<br>15:56 Analog AI <br>36:08 The strategic shift of Subnet 46 <br>47:55 Hardware innovations<br>01:02:09 Chain… <a href="https://t.co/6oDUXClQIg">pic.twitter.com/6oDUXClQIg</a></p>&mdash; Mark Jeffrey (@markjeffrey) <a href="https://x.com/markjeffrey/status/2087338149978726804?ref_src=twsrc%5Etfw">August 12, 2026</a></blockquote><script async src="https://platform.x.com/widgets.js" charset="utf-8"></script></div> </div></figure> <p class="wp-block-paragraph">The subnet&#8217;s founder appeared on<a href="https://taodaily.io/hash-rate-ep-164-beyond-sports-betting-djinns-marketplace-for-verifiable-intelligence/"> Hash Rate Pod</a> to explain Managed Virtual Network Integration Runtime (Managed VNet IR), a pattern he originally built at Microsoft, where his own code tunnels directly into miner infrastructure and runs the inference itself.</p> <p class="wp-block-paragraph">Miners bring the hardware, Instant brings the software, and that inversion is why the subnet already produces 300 to 500 tokens per second on GPT-OSS 20B (roughly double what OpenRouter delivers) with a roadmap stretching toward model weights baked into analog silicon.</p> <h2 class="wp-block-heading">Key Insights From the Episode</h2> <p class="wp-block-paragraph">The episode moved through the pivot from Zipcode to Instant, the Ditto customer relationship, the throughput benchmarks, the Managed VNet IR pattern, and the analog computing endgame.</p> <p class="wp-block-paragraph">1. <strong>Instant is SN46&#8217;s second life:</strong> The subnet started as <a href="https://x.com/instantsubnet/status/2085447603404087381?s=20">Zipcode</a>, building the ZipUSD stablecoin and a plan to bring real estate loans on-chain.</p> <p class="wp-block-paragraph"><a href="https://taodaily.io/bittensor-v440-restructures-emissions-to-favor-demand-over-inactivity/">V440</a> shortened the runway required for that plan, so Zipcode moved to Base and the subnet slot pivoted into inference.</p> <figure class="wp-block-embed is-type-rich is-provider-x wp-block-embed-x"><div class="wp-block-embed__wrapper"> <div class="embed-x"><blockquote class="twitter-tweet" data-width="500" data-dnt="true"><p lang="en" dir="ltr">Yesterday we announced the pivot to <a href="https://x.com/instantsubnet?ref_src=twsrc%5Etfw">@instantsubnet</a> on Subnet 46!<br><br>To confirm, the team does not change. The brand and revenue model changes.<br><br>We firmly believe this is best for Stakers of 46 and allows our team to shine!<br><br>Onwards and upwards, as always. 🧵 <a href="https://t.co/GFTkv4o9to">https://t.co/GFTkv4o9to</a></p>&mdash; Zipcode (@zipcodenetwork) <a href="https://x.com/zipcodenetwork/status/2085800585274531988?ref_src=twsrc%5Etfw">August 7, 2026</a></blockquote><script async src="https://platform.x.com/widgets.js" charset="utf-8"></script></div> </div></figure> <p class="wp-block-paragraph">2. <a href="https://taodaily.io/dittos-two-subnet-architecture-to-combine-persistent-memory-with-high-speed-inference/"><strong>Ditto (SN118)</strong></a><strong> is the first customer, and the deal is real:</strong> Ditto ran into throughput ceilings first on Chutes, then on OpenRouter, but they needed faster iteration than either could provide. Instant goes live for Ditto on Monday, dialing up through the week rather than flipping a hard switch.</p> <figure class="wp-block-embed is-type-rich is-provider-x wp-block-embed-x"><div class="wp-block-embed__wrapper"> <div class="embed-x"><blockquote class="twitter-tweet" data-width="500" data-dnt="true"><p lang="en" dir="ltr">Introducing Instant!<br><br>Lightning fast private inference, on Subnet 46.<br><br>Our first customer <a href="https://x.com/heydittoai?ref_src=twsrc%5Etfw">@heydittoai</a> has signed agreements to access Instant for DittoBench.<br><br>Faster inference = Faster innovation <a href="https://t.co/DpozwZexYW">pic.twitter.com/DpozwZexYW</a></p>&mdash; Instant Inference (@instantsubnet) <a href="https://x.com/instantsubnet/status/2085447603404087381?ref_src=twsrc%5Etfw">August 6, 2026</a></blockquote><script async src="https://platform.x.com/widgets.js" charset="utf-8"></script></div> </div></figure> <p class="wp-block-paragraph">3. <strong>Current throughput (300 to 500 tokens per second on GPT-OSS 20B):</strong> OpenRouter sits around 253 tokens per second on the same model today. Instant has already pushed internal builds past 500 in personal testing, which is roughly double the incumbent throughput.</p> <p class="wp-block-paragraph">4. <strong>Managed VNet IR is the trick nobody else on Bittensor is using:</strong> He built the pattern originally at Microsoft on Purview when Royal Bank of Canada could not legally allow data ingestion. Miners download a package, open a tunnel into their hardware, and Instant&#8217;s code runs the inference directly. No TEE attestation required because the code is his rather than theirs.</p> <figure class="wp-block-image size-large"><img fetchpriority="high" decoding="async" width="1024" height="465" src="https://taodaily.io/wp-content/uploads/2026/08/image-87-1024x465.png" alt="" class="wp-image-23883" srcset="https://taodaily.io/wp-content/uploads/2026/08/image-87-1024x465.png 1024w, https://taodaily.io/wp-content/uploads/2026/08/image-87-300x136.png 300w, https://taodaily.io/wp-content/uploads/2026/08/image-87-768x348.png 768w, https://taodaily.io/wp-content/uploads/2026/08/image-87.png 1880w" sizes="(max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption"><a href="https://instantsubnet.com/">Join Instant’s Waitlist</a></figcaption></figure> <p class="wp-block-paragraph">5. <strong>That pattern unlocks a much bigger hardware roadmap:</strong> Miners can bring any hardware profile they want, from consumer GPUs to FPGAs to eventually analog silicon. The mechanism does not care what runs the inference, only how many tokens per second come back.</p> <p class="wp-block-paragraph">6. <strong>The endgame is analog computing chips:</strong> Taalas (just acquired by AMD) runs Llama 3.1 8B on a 6nm chip. Mythic runs YOLO V8 at 30 frames per second on 10 watts and Llama 3 7B at 26,000 tokens per second per user.&nbsp;</p> <figure class="wp-block-image size-large"><img decoding="async" width="1024" height="232" src="https://taodaily.io/wp-content/uploads/2026/08/image-86-1024x232.png" alt="" class="wp-image-23882" srcset="https://taodaily.io/wp-content/uploads/2026/08/image-86-1024x232.png 1024w, https://taodaily.io/wp-content/uploads/2026/08/image-86-300x68.png 300w, https://taodaily.io/wp-content/uploads/2026/08/image-86-768x174.png 768w, https://taodaily.io/wp-content/uploads/2026/08/image-86.png 1577w" sizes="(max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption"><a href="https://ir.amd.com/news-events/press-releases/detail/1296/amd-acquires-taalas-to-advance-compute-solutions-for-rapidly-growing-ai-inference-market">AMD Acquires Taalas</a></figcaption></figure> <p class="wp-block-paragraph">7. <strong>The vision is Fable running on an Apple Watch:</strong> Model weights baked into resistors, inference happening at the speed of physics, zero translation laye