← Retour au fil
IOTA : l'expérience Orion-100B prouve que l'entraînement de modèles de pointe n'exige pas un seul immense centre de données
TAO Daily26 sept., 12h · il y a 5j

IOTA : l'expérience Orion-100B prouve que l'entraînement de modèles de pointe n'exige pas un seul immense centre de données

Entraîner un modèle de 100 milliards de paramètres sans data center géant : Macrocosmos l'a fait avec 48 GPU A100 reliés par simple internet, à 30,8% de MFU moyenne.

L'équipe Macrocosmos, à l'origine de Bittensor Subnet 9 (IOTA/SN9), a entraîné un modèle de 100 milliards de paramètres issu d'une architecture Llama-3.2 modifiée sur 48 GPU A100-80GB répartis dans cinq data centers américains, reliés par internet classique. Résultat : 30,8% de MFU moyenne, pics soutenus à 38%, environ 9 000 tokens par seconde et 65% de la performance d'une configuration équivalente regroupée au même endroit.

L'exploit repose sur le parallélisme en pipeline (16 étapes, 3 réplicas), la compression d'activations ResBM, un réseau pair-à-pair tolérant aux pannes et une synchronisation distribuée. Arrêté après 1,1 milliard de tokens traités pour raisons de coût, l'expérience revendique le meilleur MFU rapporté pour un entraînement distribué en pipeline — un signal fort pour l'IA décentralisée.

Bittensor

Détails

Source
TAO Daily
Publication
26 sept. à 12h23

Contenu source (brut)

<p class="wp-block-paragraph">What if training a 100-billion-parameter AI model didn&#8217;t require putting hundreds of GPUs in one highly specialised data centre?</p> <p class="wp-block-paragraph">That is the question behind <strong>IOTA’s Orion-100B experiment</strong>.</p> <p class="wp-block-paragraph">On 24 September, IOTA (SN9) brought attention back to <a href="https://taodaily.io/train-at-home-expands-to-16-countries-as-iotas-distributed-training-swarm-goes-global/">the June 2026 run</a>, in which Macrocosmos trained a 100-billion-parameter model across <strong>five US data centres</strong>, using individual A100 GPUs connected through ordinary internet infrastructure.</p> <p class="wp-block-paragraph">The experiment reached <strong>30.8% average Model FLOPs Utilization (MFU)</strong>, with a sustained peak of <strong>38%</strong>, while achieving roughly <strong>65% of the performance of an equivalent co-located setup</strong>.</p> <p class="wp-block-paragraph">It processed around <strong>1.1 billion tokens over two days</strong> before being stopped because of cost.</p> <p class="wp-block-paragraph">This is significant because the training was done across geographically separated machines that were never connected by the kind of specialised networking normally associated with large-scale AI training.</p> <figure class="wp-block-embed is-type-rich is-provider-x wp-block-embed-x"><div class="wp-block-embed__wrapper"> <div class="embed-x"><blockquote class="twitter-tweet" data-width="500" data-dnt="true"><p lang="en" dir="ltr">We trained a 100B parameter model across five data centres, on single A100s talking to each other over ordinary internet links.<br><br>It held 30.8% average MFU, with a sustained 38% peak, and ran at roughly 65% of the speed the same job reaches on co-located high-bandwidth hardware.…</p>&mdash; IOTA ・ SN9 (@IOTA_SN9) <a href="https://x.com/IOTA_SN9/status/2103175140699934764?ref_src=twsrc%5Etfw">September 24, 2026</a></blockquote><script async src="https://platform.x.com/widgets.js" charset="utf-8"></script></div> </div></figure> <h2 class="wp-block-heading"><strong>What Orion-100B Did</strong></h2> <figure class="wp-block-image size-large is-resized"><img loading="lazy" decoding="async" width="1024" height="416" src="https://taodaily.io/wp-content/uploads/2026/09/image-4-1024x416.jpeg" alt="" class="wp-image-25052" style="aspect-ratio:2.302583025830258;width:624px;height:auto" srcset="https://taodaily.io/wp-content/uploads/2026/09/image-4-1024x416.jpeg 1024w, https://taodaily.io/wp-content/uploads/2026/09/image-4-300x122.jpeg 300w, https://taodaily.io/wp-content/uploads/2026/09/image-4-768x312.jpeg 768w, https://taodaily.io/wp-content/uploads/2026/09/image-4.jpeg 2048w" sizes="(max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">IOTA scales to more GPUs, and converges more reliably</figcaption></figure> <p class="wp-block-paragraph">Macrocosmos, the team behind Bittensor Subnet 9 and IOTA, trained a modified Llama-3.2 architecture scaled to <strong>100 billion parameters</strong>.</p> <p class="wp-block-paragraph">The model was divided across <strong>16 pipeline-parallel stages</strong>, with three replicas, using a total of <strong>48 single A100-80GB GPUs</strong>.</p> <p class="wp-block-paragraph">Those GPUs were spread across <strong>five US data centres</strong> and communicated over commodity internet connections rather than a specialised high-bandwidth fabric.</p> <p class="wp-block-paragraph">The reported results were:</p> <ul class="wp-block-list"> <li><strong>30.8% average MFU</strong></li> <li><strong>38% sustained peak MFU</strong> over six hours</li> <li>Around <strong>65% of the performance</strong> of the equivalent co-located setup</li> <li>Approximately <strong>9,000 tokens per second</strong> on average</li> <li>Around <strong>1.1 billion tokens processed</strong> before the run was stopped</li> </ul> <p class="wp-block-paragraph">Macrocosmos described the result as the highest MFU reported for a distributed pipeline-parallel run. The TAO Daily also covered Orion-100B at the time as a major milestone for decentralised AI training.</p> <figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="804" src="https://taodaily.io/wp-content/uploads/2026/09/image-259-1024x804.png" alt="" class="wp-image-25053" srcset="https://taodaily.io/wp-content/uploads/2026/09/image-259-1024x804.png 1024w, https://taodaily.io/wp-content/uploads/2026/09/image-259-300x236.png 300w, https://taodaily.io/wp-content/uploads/2026/09/image-259-767x602.png 767w, https://taodaily.io/wp-content/uploads/2026/09/image-259.png 1182w" sizes="(max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">The source is an <a href="https://x.com/taodaily_io/status/2062601833596747979">X article</a>.</figcaption></figure> <h2 class="wp-block-heading"><strong>Why Training Over the Internet Is Difficult</strong></h2> <p class="wp-block-paragraph">Training a model this large is not as simple as connecting a collection of GPUs and pressing start.</p> <p class="wp-block-paragraph">In a conventional AI cluster, GPUs can communicate over extremely fast networking. Orion had to work with machines separated across different locations and connected through ordinary internet infrastructure.</p> <p class="wp-block-paragraph">IOTA addressed this by using <strong>pipeline parallelism</strong>.</p> <p class="wp-block-paragraph">Instead of requiring every GPU to handle the entire model, the model is divided into stages. Each group of GPUs works on a different part of the model, allowing the workload to be distributed across the network.</p> <p class="wp-block-paragraph">The problem is then keeping those stages supplied with data and synchronised despite the distance between machines.</p> <p class="wp-block-paragraph">Three parts of IOTA&#8217;s architecture were particularly important:</p> <ul class="wp-block-list"> <li><strong>ResBM (Residual Bottleneck Model):</strong> a lossless activation-compression technique designed to reduce the amount of data transferred between machines.</li> <li><strong>A fault-tolerant peer-to-peer networking protocol:</strong> designed to keep training operating despite unreliable connections or individual failures.</li> <li><strong>Distributed synchronisation:</strong> used to keep the different parts of the training process coordinated.</li> </ul> <p class="wp-block-paragraph">The technology was developed through earlier experimentation, including a <strong>1.5-billion-parameter testbed</strong> that was used for hundreds of controlled experiments before the team moved toward larger runs.</p> <figure class="wp-block-embed is-type-wp-embed is-provider-the-tao-daily wp-block-embed-the-tao-daily"><div class="wp-block-embed__wrapper"> <blockquote class="wp-embedded-content" data-secret="22dlCECpkG"><a href="https://taodaily.io/macrocosmos-just-cracked-one-of-decentralized-ais-hardest-problems/">Macrocosmos Just Cracked One of Decentralized AI&#8217;s Hardest Problems</a></blockquote><iframe class="wp-embedded-content" sandbox="allow-scripts" security="restricted" style="position: absolute; visibility: hidden;" title="“Macrocosmos Just Cracked One of Decentralized AI’s Hardest Problems” — The TAO Daily" src="https://taodaily.io/macrocosmos-just-cracked-one-of-decentralized-ais-hardest-problems/embed/#?secret=VQf3nFx9D4#?secret=22dlCECpkG" data-secret="22dlCECpkG" width="500" height="282" frameborder="0" marginwidth="0" marginheight="0" scrolling="no"></iframe> </div></figure> <h2 class="wp-block-heading"><strong>The Economics Matter Too</strong></h2> <p class="wp-block-paragraph">There was another reason Orion-100B was significant: <strong>the hardware did not need to be packaged into one expensive machine.</strong></p> <p class="wp-block-paragraph">Mac