# Welcome to Kinesis Network

Kinesis Network is a unified compute platform built to simplify how organizations run and scale modern workloads. From AI and machine learning to data processing and high-performance applications, Kinesis provides the flexibility to access the right compute, at the right time, without the operational overhead of traditional infrastructure.

By combining Kinesis-managed resources with bring-your-own compute (BYOC), the platform enables teams to operate across environments with consistency and control. A container-based runtime ensures portability, while a unified interface provides clear visibility into performance and cost, helping teams make informed decisions as workloads evolve.

Kinesis is designed around efficiency and transparency. With usage-based pricing for on-demand workloads and support for pre-provisioned compute, organizations can optimize for both agility and predictability — without compromising on performance.


# What You Can Do with Kinesis

**Run AI and LLM workloads**\
Deploy, fine-tune, and serve machine learning models using high-performance GPU infrastructure, with the flexibility to scale as demand changes.

**Build and operate data pipelines**\
Execute batch and streaming workloads efficiently, with the ability to adapt compute resources to changing data volumes.

**Optimize cost and utilization**\
Choose between on-demand, usage-based compute or pre-configured resources, and gain a unified view of performance and spend across your workloads.

**Leverage existing infrastructure**\
Connect your own compute resources to Kinesis and manage them alongside platform-provided capacity, all within a single control plane.

**Standardize your runtime environment**\
Use container-based deployments to ensure consistency across development, testing, and production environments.


# Our Core Tenets

At Kinesis Network, we are guided by principles that ensure our customers never have to compromise. These tenets drive everything we build and innovate:

1. **Reliability**
   * Our technology is designed to be dependable because we know our customers’ businesses—and their own customers—rely on us. Reliability is a foundation we never compromise on.
2. **Security, Privacy, and Confidentiality**
   * We adhere to strict compliance standards and only use resources that meet our rigorous benchmarks to safeguard customer data and ensure confidentiality.
3. **Performance**
   * We deliver optimal performance for all workloads, unless customers intentionally choose to trade performance for cost savings.
4. **Transparency**
   * We operate with complete transparency, never using tools or practices that we cannot proudly explain to our customers.
5. **WYSIWYG (What You See Is What You Get)**
   * With Kinesis, there are no hidden costs or surprises. Our customers always know exactly what they are paying for.
6. **Simplicity**
   * Simplicity is at the core of our product. We hide backend complexities behind an intuitive, user-friendly interface, making everything—from usage to pricing—straightforward and easy to navigate.

These principles are the backbone of Kinesis Network, ensuring trust, clarity, and excellence in every aspect of our platform.


# How Does Kinesis Network Enhance Customer Value with Optimized Compute Service

As computing demands surge, particularly in AI and machine learning, enterprises require scalable, high-performance infrastructure. Kinesis Network meets this need with a serverless, fully managed Compute Cloud that leverages pooled, unused resources to deliver efficient and sustainable compute power. This post explores how Kinesis Network provides enterprise customers with the tools they need to focus on innovation, while simplifying and reducing the costs of infrastructure management.

**Key Challenges in Compute Infrastructure**

1. **Rising Costs**: AI-driven organizations often face high expenses related to GPU resources. With Kinesis Network’s pay-as-you-go model, enterprises pay only for active compute usage, avoiding the idle costs common with traditional cloud services.
2. **Sustainable Resource Use**: Many businesses are concerned with the environmental impact of underutilized compute resources. By pooling idle capacity, Kinesis Network helps reduce waste, contributing to a lower carbon footprint.
3. **Complex Infrastructure Management**: Not all companies have the specialized skills to handle GPU-intensive infrastructures. Kinesis Network addresses this by offering a fully managed platform that eliminates infrastructure complexities, allowing companies to focus on their core applications.

**Proprietary Technology for Seamless Compute Sourcing**

Kinesis Network’s proprietary software optimizes compute resource allocation from diverse sources. Key features include:

* **Resource Pooling and Tunneling**: By aggregating unused GPU resources, Kinesis Network provides scalable, high-performance compute power, allowing companies to access robust infrastructure without the need for dedicated hardware.
* **Automated Scaling**: The platform dynamically allocates resources based on demand, making it suitable for a wide range of workloads, from periodic batch processing to continuous applications.
* **On-Demand Flexibility**: Kinesis Network’s on-demand scaling, paired with a per-second billing model, enables enterprises to manage variable workloads effectively, avoiding long-term commitments and unnecessary costs.

**Benefits for Enterprise Customers**

1. **Cost Efficiency**: With a usage-based pricing model, Kinesis Network reduces costs compared to traditional cloud services, allowing businesses to optimize their budgets and pay only for the resources they actively use.
2. **Scalability and Flexibility**: Access to the latest high-performance GPUs makes it easy for enterprises to handle demanding applications like AI training and inference, while allowing them to scale resources up or down as needed.
3. **Quick Onboarding**: The platform streamlines the onboarding process. Businesses can directly upload Docker containers, configure requirements, and let Kinesis Network handle resource management, helping companies deploy projects faster.
4. **Sustainability**: Kinesis Network’s approach to pooling unused capacity not only optimizes resource use but also aligns with companies’ goals to reduce their environmental impact.
5. **Multi-cloud**: Kinesis Network makes multi-cloud goals easy. Customers simply specify their desired cloud distribution—for example, 60% AWS, 30% Azure, and 10% GCP—and we handle the rest. Our unified compute pools across major providers allow seamless VM selection, giving you flexibility and control.

Kinesis Network is an ideal solution for enterprises seeking to improve efficiency, scalability, and sustainability in their compute infrastructure. Its fully managed, serverless model allows companies to focus on innovation, reducing costs and minimizing the complexities of infrastructure management. For customers ready to optimize their infrastructure, Kinesis Network provides a compelling, forward-thinking solution.


# How does Kinesis Network Work?

Kinesis Network operates with the efficiency, ease of use, and affordability of a power outlet—but offers greater control and transparency.

The electricity in your home comes from diverse sources—large power plants (nuclear, solar, wind, hydro) and smaller generators like your neighbors' rooftop solar panels. The power grid manages and conditions this electricity, delivering it to you in a consolidated way. You receive a single bill, regardless of the source. There's no need to hesitate before using power; you can host a dinner party or run the dishwasher without worry. The supply of electricity scales invisibly with your demand. If you choose, you can even install your own solar panels and sell excess power back to the grid, lowering your bills.

**Kinesis Network assembles a diverse array of compute participants, varying in type and scale, from around the world into a unified compute grid.**

Kinesis' compute providers span a wide spectrum, ranging from large data centers to smaller facilities and even individual devices like gaming computers in homes or idle machines in offices.

As a customer, you maintain complete control over your compute source. For proof-of-concept work, the most cost-effective option might be using compute resources from around the world. However, for sensitive tasks involving personally identifiable information (PII) or subject to regulatory compliance, dedicated data centers may be more appropriate. Additionally, if your existing infrastructure is already based in a specific data center (such as AWS or Azure), staying within that ecosystem could be optimal.

**Kinesis Network consolidates the compute**

Although our network allows low-level access to individual resources in the underlying infra, the real value comes from seamless unification of those resources in an easy to use 1-2-3 fashion.

1. Customers upload their containers (in Docker or other supported formats) through the Kinesis Portal.
2. Afterwards, they make their preferences in terms of minimum hardware specifications (such as minimum RAM, VRAM, CPU speed needed).
3. Upon pressing the start button, the Portal provisions the best servers from the grid. It constantly monitors the performance of the group of servers, and replaces them as needed, or adds/removes to/from the groups compute power.

These happens entirely seamlessly, automatically and responsively.

**Kinesis Network gives secure and easy access to your workloads**

Your applications can easily and securely access containers running in the grid. While direct HTTP or raw TCP access to individual servers is available (e.g. for development purposes), we provide a front-end load balancer to simplify the backend complexity. When a customer makes a request, the nearest data center responds with the most performant server available at the moment.

Individual resources are instantly scaled up, down, in, or out as needed. If the health of some servers degrades, new ones are activated. This orchestration is seamless and automatic, operating within your specified preferences and performance goals.

**You pay for only what you actually use**

Similar to your electric bill, you only pay for what you actually use, not for what you've reserved or could have used. We meticulously track every instance of usage—down to milliseconds for CPU/GPU usage and bytes for network I/O, storage, and RAM/VRAM occupation. Despite the intricate metering and billing processes, we simplify it for you: you receive a single, comprehensive bill.

For a deeper understanding of your usage, we provide detailed dashboards that show where and when activity occurs. Unlike instance or VM-based workloads, our dashboards display usage per application, revealing the true sources of costs. This insight allows you to target specific areas for optimization and better inform you about your business needs.

**Kinesis Network is economical**

There is plenty of compute in the world, but

* some is overly expensive — mega datacenters can charge high premiums
* some is not easily accessible — despite idle capacity in data centers or homes, it remains untapped
* and some is inefficiently utilized — cloud providers often offer fixed-size instances, leading to wasted resources like unused RAM, VRAM, or CPU power. For example, virtual machines typically use less than 20% of their CPU capacity, yet still paying for 100% of the reservation.

Kinesis Network mobilizes underutilized computing resources, passing the savings on to you.

**You can contribute to the network in exchange for credits or the specific compute resources you need.**

If you have idle compute resources (perhaps due to over-provisioned reserved instances from AWS, Azure, or other providers), you can offer them back to the grid. As these resources contribute to the network, you'll earn credits for your own use—or even turn a profit if you contribute more than you consume. This system also allows you to convert excess capacity from one type to another. For instance, if you have surplus CPU power but need GPU capacity for an upcoming AI applications, you can efficiently trade your CPU resources on the Kinesis Network for the GPU power you currently require.


# Running LLMs with Kinesis Network

Running large language models shouldn’t require navigating fragmented infrastructure, overpaying for idle capacity, or locking into a single cloud.

**Kinesis Network** is a unified compute platform designed to simplify and optimize how LLM workloads are deployed, scaled, and managed. It enables developers and teams to access high-performance GPUs on demand, with transparent, True-Util pricing that charges only for what you actually use. Whether you’re fine-tuning models, running inference at scale, or experimenting with new architectures, Kinesis provides a consistent, container-based environment that removes operational friction while maximizing performance, flexibility, and cost efficiency.


# Creating a project

Projects in Kinesis Network are the foundational unit for organizing your workloads. A project represents a logical grouping of applications that belong together, whether that’s a single experiment, a production deployment, or a team’s shared environment.

By grouping apps into a project, you gain a unified view of both performance and cost across all related workloads. This makes it easier to monitor resource utilization, track spend, and understand how different components of your LLM stack behave collectively. Instead of managing each application in isolation, projects provide a structured way to operate, optimize, and scale your workloads with clarity.

Please go to <https://portal.kinesis.network/dashboard/projects> to create your first project.

<figure><img src="/files/eWy6VIoOdh9x0frYrjiq" alt=""><figcaption></figcaption></figure>


# Creating Apps with App Gallery

When you create a project, App creation menus are shown automatically. You can also create a new app by clicking Create App button:

<figure><img src="/files/ujUsKz4CRw6hTzU6fMj2" alt=""><figcaption></figcaption></figure>

Once you are inside the app creation dialog, you have two options:

* You can quickly create an LLM app by choosing from our pre-curated library.
* You can create a custom app by using code or Docker images.

### Creating an App with App Gallery

Create apps quickly by selecting from the hand-curated App Gallery. Access it by clicking “Select Template from Gallery” at the top or choosing the App Gallery menu on the left.

<figure><img src="/files/kDaeKxCXhqQTBoGCMqHz" alt=""><figcaption></figcaption></figure>

Once in the App Gallery, you can choose from many options. Please don’t miss the first one that offers a very large selection of latest and greatest LLM models.

<figure><img src="/files/xFHDs8LgZpAuZ7GfSz4i" alt=""><figcaption></figcaption></figure>

Once you make your selection, you have the option to customize its settings. For most purposes, “Quick Launch” will be the best. However, if you need to configure the Environment Variables for more advanced authentication and other settings, choose “Customize First”.

<figure><img src="/files/stHRNsphYOWLqXKmJ5Qv" alt=""><figcaption></figcaption></figure>

This will bring you back to App Creation dialog. As the final step, you will need to choose your compute source. Please see [Understanding Compute Sources](/getting-started/running-llms-with-kinesis-network/understanding-compute-sources)


# Creating an App from Scratch

App creation has 5 basic steps.

**Source**: This is where you connect&#x20;

* Your code from **GitHub**, or&#x20;
* Choose an image from **DockerHub**, or&#x20;
* Connect your private repository with **Docker Link**

**Runtime:** This is where you set environment variables, restart policies, health checks and more.

**Network:** This is where you configure load balancers, custom domains and firewall settings.

**Hardware:** This is where you define the minimum requirements for your application in terms of hardware resources.

**Location:** This is where you choose where to run your applications in terms of type of compute.

<figure><img src="/files/9ehlIWI2IDOuUtUGZhAZ" alt=""><figcaption></figcaption></figure>


# Understanding Compute Sources

{% hint style="info" %}
Skip this section if your administrator shares Grids with your account.
{% endhint %}

**Shared Compute: Kinesis-managed, multi-tenant, pay-per-use** The primary benefit is that applications enjoy full access to the machine's capabilities, while costs are incurred only for actual usage, measured by:

* CPU milliseconds consumed
* GPU milliseconds consumed
* RAM and VRAM bytes used

When an application is idle and not utilizing CPU or GPU resources, no corresponding charges apply. This mode is well-suited for rapid prototyping and non-critical workloads where consistent performance is not essential.

<figure><img src="/files/UWmsHYwoP1IzKSzphDOK" alt=""><figcaption></figcaption></figure>

**Grids: Compute you control—either provisioned by Kinesis or connected from your own infrastructure. Can be shared across apps or accounts.**

Grids consist of dedicated resources managed by you or your company’s system administrators. These resources are shared only within the same organization. They are billed based on the time they are reserved. Charges occur as follows:

* **Full price when active:** Servers are considered active when they have one or more apps that are not in the “Stopped” state (e.g. they are “Running” or “Starting”)
* **Storage-only price when hibernated:** If a server has no apps on it, or if all apps are in the stopped state, Kinesis hibernates it to save on costs. In this case, charges apply only to storage.

💡An app can access multiple pools, expanding its resource options.

{% hint style="info" %}
Use your organization’s existing grids when available. Confirm the appropriate pool with your system administrator. Please see next section on how to use Grids assigned to you.
{% endhint %}

<figure><img src="/files/cM7YRdoBSIFsOEMKKAwU" alt=""><figcaption></figcaption></figure>


# Understanding Grids

Grids are logical grouping of compute resources. You can target your apps on one or more grids, combining their compute power flexibly.&#x20;

If you have grids, you can share with others. Similarly, others (or your admin) can share grids with you for your perusal.&#x20;

Below is how you can choose one or more grids when configuring and launching your apps.&#x20;

<figure><img src="/files/bmJOWEq8vQ3lOa12SwDl" alt=""><figcaption></figcaption></figure>


# Monitoring Apps

Once you launch your app, it will take a minute or so to provision it and run it in a container. Your app is accessible through the endpoints displayed on the upper right:

<figure><img src="/files/RhBNAUi1zQSjK9KxN2lc" alt=""><figcaption></figcaption></figure>

Just below, you will several tabs, allowing you to monitor and interact with your app.

### Logs

Real-time logs generated by your application will appear in this section. For applications utilizing multiple servers, you have the capability to view logs from each server separately.

<figure><img src="/files/WaZWN1Y1BdSKFc20jyAk" alt=""><figcaption></figcaption></figure>

### Terminal

SSH into your container for advanced settings and troubleshooting. Access each server individually by selecting the relevant tab.

<figure><img src="/files/W5aFqRYQiXrXD6zXyrUS" alt=""><figcaption></figcaption></figure>

### Metrics

View your app's performance live or over custom time periods.

<figure><img src="/files/lwx48bUTfiY7MtFHcyBf" alt=""><figcaption></figcaption></figure>


# Support

For support, please contact

* email: <support@kinesis.network>
* or join our Discord server: <https://discord.gg/ur9g7EctPt>


# PyTorch with WireGuard

How Kinesis supports PyTorch DDP?

By Toshihito Kikuchi

[PyTorch DDP](https://docs.pytorch.org/tutorials/intermediate/ddp_tutorial.html) (= Distributed Data Parallel) is becoming the industry standard to parallelize your model training process across multiple GPUs and machines. Since one of our missions is to provide scalable GPU computes, it’s essential to support PyTorch DDP on Kinesis Network.

Does Kinesis Network support PyTorch DDP? Yes, but it isn't quite "plug-and-play" yet — it requires a bit of manual tuning to get everything running smoothly. We will fully integrate it soon, but for now, you need some special configuration to run it. In this article, I’d like to explain why it’s a little bit tricky and what we're implementing behind the scene.

### Challenge in PyTorch networking

Okay, I know you have a model to train. You wrote a python script with [DistributedDataParallel](https://docs.pytorch.org/docs/stable/generated/torch.nn.parallel.DistributedDataParallel.html), something like this.

```python
import os
import torch
import torch.nn as nn
import torch.optim as optim
import torch.distributed as dist
from torch.nn.parallel import DistributedDataParallel as DDP

def setup():
    # Initialize the process group
    # NCCL is the standard backend for NVIDIA GPUs
    dist.init_process_group(backend="nccl")
    torch.cuda.set_device(int(os.environ["LOCAL_RANK"]))

def cleanup():
    dist.destroy_process_group()

def run_training():
    setup()

    local_rank = int(os.environ["LOCAL_RANK"])
    steps = int(os.environ["STEPS"])
    rank = int(os.environ["RANK"])
    device = torch.device(f"cuda:{local_rank}")

    # 1. Define a tiny model
    model = nn.Linear(10, 10).to(device)
    model = DDP(model, device_ids=[local_rank])

    # 2. Setup Loss and Optimizer
    loss_fn = nn.MSELoss()
    optimizer = optim.SGD(model.parameters(), lr=0.001)

    # 3. Simple training loop (Synthetic data)
    print(f"[Rank {rank}] Starting training...")

    for step in range(steps):
        # Create random data on the fly
        inputs = torch.randn(20, 10).to(device)
        labels = torch.randn(20, 10).to(device)

        optimizer.zero_grad()
        outputs = model(inputs)
        loss = loss_fn(outputs, labels)
        loss.backward()
        optimizer.step()

        if step % 10 == 0 and rank == 0:
            print(f"Step {step} | Loss: {loss.item():.4f}")

    print(f"[Rank {rank}] Training complete.")
    cleanup()

if __name__ == "__main__":
    run_training()
```

We usually use `torchrun` to kick a distributed training process.  Below is the output to run a single-node, single-gpu training inside a container with detailed logs.

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 torchrun \
  --nnodes=1 --nproc_per_node=1 --node_rank=0 \
  --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=127.0.0.1:29500 \
  train.py
I0406 17:13:28.244000 1265 torch/distributed/run.py:735] Using nproc_per_node=1.
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 127.0.0.1:29500
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'timeout': 900}
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_z0qxtusm
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 17:13:28.340000 1265 torch/distributed/launcher/api.py:224]
I0406 17:13:28.344000 1265 torch/distributed/elastic/agent/server/api.py:898] [default] starting workers for entrypoint: python3
I0406 17:13:28.345000 1265 torch/distributed/elastic/agent/server/api.py:693] [default] Rendezvous'ing worker group
I0406 17:13:28.481000 1265 torch/distributed/elastic/agent/server/api.py:539] [default] Rendezvous complete for workers. Result:
I0406 17:13:28.481000 1265 torch/distributed/elastic/agent/server/api.py:539]   restart_count=0
I0406 17:13:28.481000 1265 torch/distributed/elastic/agent/server/api.py:539]   master_addr=2f3d44d8b0ea
I0406 17:13:28.481000 1265 torch/distributed/elastic/agent/server/api.py:539]   master_port=43093
I0406 17:13:28.481000 1265 torch/distributed/elastic/agent/server/api.py:539]   group_rank=0
I0406 17:13:28.481000 1265 torch/distributed/elastic/agent/server/api.py:539]   group_world_size=1
I0406 17:13:28.481000 1265 torch/distributed/elastic/agent/server/api.py:539]   local_ranks=[0]
I0406 17:13:28.481000 1265 torch/distributed/elastic/agent/server/api.py:539]   role_ranks=[0]
I0406 17:13:28.481000 1265 torch/distributed/elastic/agent/server/api.py:539]   global_ranks=[0]
I0406 17:13:28.481000 1265 torch/distributed/elastic/agent/server/api.py:539]   role_world_sizes=[1]
I0406 17:13:28.481000 1265 torch/distributed/elastic/agent/server/api.py:539]   global_world_sizes=[1]
I0406 17:13:28.481000 1265 torch/distributed/elastic/agent/server/api.py:539]   event_log_handler=null
I0406 17:13:28.481000 1265 torch/distributed/elastic/agent/server/api.py:539]
I0406 17:13:28.482000 1265 torch/distributed/elastic/agent/server/api.py:701] [default] Starting worker group
I0406 17:13:28.482000 1265 torch/distributed/elastic/agent/server/local_elastic_agent.py:299] use_agent_store: True
I0406 17:13:28.482000 1265 torch/distributed/elastic/agent/server/local_elastic_agent.py:195] Environment variable 'TORCHELASTIC_ENABLE_FILE_TIMER' not found. Do not start FileTimerServer.
I0406 17:13:28.482000 1265 torch/distributed/elastic/agent/server/local_elastic_agent.py:239] Environment variable 'TORCHELASTIC_HEALTH_CHECK_PORT' not found. Do not start health check.
2f3d44d8b0ea:1297:1297 [0] NCCL INFO ENV/Plugin: Could not find: libnccl-env.so
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Bootstrap: Using eth0:172.17.0.2<0>
2f3d44d8b0ea:1297:1297 [0] NCCL INFO cudaDriverVersion 12080
2f3d44d8b0ea:1297:1297 [0] NCCL INFO NCCL version 2.28.9+cuda12.9
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Comm config Blocking set to 1
2f3d44d8b0ea:1297:1297 [0] NCCL INFO NET/Plugin: Could not find: libnccl-net.so
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Failed to open libibverbs.so[.1]
2f3d44d8b0ea:1297:1297 [0] NCCL INFO transport/net_ib.cc:852 -> 3
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Failed to initialize NET plugin IB
2f3d44d8b0ea:1297:1297 [0] NCCL INFO NET/Socket : Using [0]eth0:172.17.0.2<0> [1]wg0:10.100.10.1<0>
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Initialized NET plugin Socket
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Assigned NET plugin Socket to comm
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Using network Socket
2f3d44d8b0ea:1297:1297 [0] NCCL INFO ncclCommInitRankConfig comm 0x3cd2a790 rank 0 nranks 1 cudaDev 0 nvmlDev 0 busId 70 commId 0xbda3ad566558fee7 - Init START
2f3d44d8b0ea:1297:1297 [0] NCCL INFO RAS client listening socket at ::1<28028>
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Bootstrap timings total 0.001224 (create 0.000070, send 0.000246, recv 0.000273, ring 0.000001, delay 0.000000)
2f3d44d8b0ea:1297:1297 [0] NCCL INFO NCCL_IGNORE_DISABLED_P2P set by environment to 1.
2f3d44d8b0ea:1297:1297 [0] NCCL INFO ncclTopoGetCpuAffinity: Affinity for GPU 0 is empty, ignoring. (GPU affinity =  ; CPU affinity = 0-27).
2f3d44d8b0ea:1297:1297 [0] NCCL INFO comm 0x3cd2a790 rank 0 nRanks 1 nNodes 1 localRanks 1 localRank 0 MNNVL 0
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Channel 00/64 : 0
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Channel 01/64 : 0
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Channel 02/64 : 0
...snip...
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Channel 63/64 : 0
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Trees [0] -1/-1/-1->0->-1 [1] -1/-1/-1->0->-1 [2] -1/-1/-1->0->-1 [3] -1/-1/-1->0->-1 [4] -1/-1/-1->0->-1 [5] -1/-1/-1->0->-1 [6] -1/-1/-1->0->-1 [7] -1/-1/-1->0->-1 [8] -1/-1/-1->0->-1 [9] -1/-1/-1->0->-1 [10] -1/-1/-1->0->-1 [11] -1/-1/-1->0->-1 [12] -1/-1/-1->0->-1 [13] -1/-1/-1->0->-1 [14] -1/-1/-1->0->-1 [15] -1/-1/-1->0->-1 [16] -1/-1/-1->0->-1 [17] -1/-1/-1->0->-1 [18] -1/-1/-1->0->-1 [19] -1/-1/-1->0->-1 [20] -1/-1/-1->0->-1 [21] -1/-1/-1->0->-1 [22] -1/-1/-1->0->-1 [23] -1/-1/-1->0->-1 [24] -1/-1/-1->0->-1 [25] -1/-1/-1->0->-1 [26] -1/-1/-1->0->-1 [27] -1/-1/-1->0->-1 [28] -1/-1/-1->0->-1 [29] -1/-1/-1->0->-1 [30] -1/-1/-1->0->-1 [31] -1/-1/-1->0->-1 [32] -1/-1/-1->0->-1 [33] -1/-1/-1->0->-1 [34] -1/-1/-1->0->-1 [35] -1/-1/-1->0->-1 [36] -1/-1/-1->0->-1 [37] -1/-1/-1->0->-1 [38] -1/-1/-1->0->-1 [39] -1/-1/-1->0->-1 [40] -1/-1/-1->0->-1 [41] -1/-1/-1->0->-1 [42] -1/-1/-1->0->-1 [43] -1/-1/-1->0->-1 [44] -1/-1/-1->0->-1 [45] -1/-1/-1->0->-1 [46] -1/-1/-1->0->-1 [47
2f3d44d8b0ea:1297:1297 [0] NCCL INFO P2P Chunksize set to 524288
2f3d44d8b0ea:1297:1297 [0] NCCL INFO PROFILER/Plugin: Could not find: libnccl-profiler.so
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Check P2P Type isAllDirectP2p 1 directMode 0 isAllCudaP2p 1
2f3d44d8b0ea:1297:1331 [0] NCCL INFO [Proxy Service] Device 0 CPU core 10
2f3d44d8b0ea:1297:1332 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 9
2f3d44d8b0ea:1297:1297 [0] NCCL INFO TUNER/Plugin: Could not find: libnccl-tuner.so
2f3d44d8b0ea:1297:1297 [0] NCCL INFO 64 coll channels, 64 collnet channels, 0 nvls channels, 64 p2p channels, 64 p2p channels per peer
2f3d44d8b0ea:1297:1297 [0] NCCL INFO CC Off, workFifoBytes 1048576
2f3d44d8b0ea:1297:1297 [0] NCCL INFO ncclCommInitRankConfig comm 0x3cd2a790 rank 0 nranks 1 cudaDev 0 nvmlDev 0 busId 70 commId 0xbda3ad566558fee7 - Init COMPLETE
2f3d44d8b0ea:1297:1297 [0] NCCL INFO Init timings - ncclCommInitRankConfig: rank 0 nranks 1 total 0.17 (kernels 0.13, alloc 0.00, bootstrap 0.00, allgathers 0.00, topo 0.00, graphs 0.00, connections 0.02, rest 0.00)
[Rank 0] Starting training...
Step 0 | Loss: 1.1732
Step 10 | Loss: 1.1904
Step 20 | Loss: 1.5663
[Rank 0] Training complete.
2f3d44d8b0ea:1297:1297 [0] NCCL INFO comm 0x3cd2a790 rank 0 nranks 1 cudaDev 0 busId 70 - Destroy COMPLETE
2f3d44d8b0ea:1297:1297 [0] NCCL INFO ENV/Plugin: Closing env plugin ncclEnvDefault
I0406 17:13:32.492000 1265 torch/distributed/elastic/agent/server/api.py:917] [default] worker group successfully finished. Waiting 300 seconds for other agents to finish.
I0406 17:13:32.493000 1265 torch/distributed/elastic/agent/server/api.py:970] Local worker group finished (WorkerState.SUCCEEDED). Waiting 300 seconds for other agents to finish
I0406 17:13:32.493000 1265 torch/distributed/elastic/agent/server/api.py:984] Done waiting for other agents. Elapsed: 0.0004515647888183594 seconds
root@2f3d44d8b0ea:/app#
```

It looks pretty straightforward. Can we just specify `--nnodes` and `--node_rank` accordingly to run this on multiple gpu nodes? Well, unfortunately no, that’s why I’m writing this article.

The challenge is networking. In the example above, I specified `--rdzv_endpoint=127.0.0.1:29500`. Does this mean each node communicates with this endpoint during a training process, like the Hub-and-Spoke topology?  The answer is no.  As the name implies, it’s a rendezvous point.  Every node meets there, to retrieve the actual endpoint to communicate with. In the output above, you see `master_addr=2f3d44d8b0ea` and `master_port=43093` , that’s the actual endpoint each node communicates with. The master port, 43093 in this case, is dynamically assigned, which means we cannot predict which port needs to be opened. This is a problem in a Docker environment because we need to expose ports on creation but we don't know it on creation.  Should we use the host network `--network=host` or publish a wide range of ports like `-p 1000:65535`?  Apparently that design is far from ideal. We want our containers to be contained as much as possible.

In [the earlier article](https://docs.kinesis.network/blog/reaching-out-to-home-computers), I wrote we leverage [WireGuard](https://www.wireguard.com/) to bring home computers. We can get help from WireGuard to solve this PyTorch situation too. Once we set up a WireGuard network on every Docker containers, they communicate with one another via a single UDP port. This journey starts here.

### WireGuard Setup

Setting up a WireGuard network is pretty easy.  I skip the detailed setup steps in this article.  Basically you need two things: 1) create a conf file like /etc/wireguard/wg0.conf and run `wg-quick up wg0`.  Besides, when you create a Docker container, you need to specify `--cap-add=NET_ADMIN` and publish a UDP port.  Lastly, you need to install the packages `wireguard-tools iptables iproute2`.  The wireguard driver exists in the host as a part of Linux kernel, but you still need client tools to use it.

Once setup is done, you will see a virtual interface like `wg0`.

```log
root@2f3d44d8b0ea:/app# ip -4 a
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc noqueue state UNKNOWN group default qlen 1000
    inet 127.0.0.1/8 scope host lo
       valid_lft forever preferred_lft forever
2: eth0@if32: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP group default  link-netnsid 0
    inet 172.17.0.2/16 brd 172.17.255.255 scope global eth0
       valid_lft forever preferred_lft forever
3: wg0: <POINTOPOINT,NOARP,UP,LOWER_UP> mtu 1420 qdisc noqueue state UNKNOWN group default qlen 1000
    inet 10.100.10.1/24 scope global wg0
       valid_lft forever preferred_lft forever
```

### Debugging with GDB

Let's run `torchrun` with the WireGuard network address 10.100.10.1 as the rendezvous point.

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 torchrun \
  --nnodes=1 --nproc_per_node=1 --node_rank=0 \
  --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500 \
  train.py
I0406 17:54:35.221000 1339 torch/distributed/run.py:735] Using nproc_per_node=1.
I0406 17:54:35.316000 1339 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'timeout': 900}
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_q6zcchrr
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 17:54:35.317000 1339 torch/distributed/launcher/api.py:224]
[E406 17:55:24.823650196 socket.cpp:1028] [c10d] The client socket has timed out after 60000ms while trying to connect to (10.100.10.1, 29500).
[W406 17:55:24.824122039 TCPStore.cpp:340] [c10d] TCP client failed to connect/validate to host 10.100.10.1:29500 - retrying (try=0, timeout=60000ms, delay=46179ms): The client socket has timed out after 60000ms while trying to connect to (10.100.10.1, 29500).
Exception raised from throwTimeoutError at /pytorch/torch/csrc/distributed/c10d/socket.cpp:1030 (most recent call first):
frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x71c56a97205d in /usr/local/lib/python3.12/dist-packages/torch/lib/libc10.so)
frame #1: <unknown function> + 0x16daa1c (0x71c4c09eea1c in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #2: <unknown function> + 0x6b25497 (0x71c4c5e39497 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #3: <unknown function> + 0x6b256cf (0x71c4c5e396cf in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #4: <unknown function> + 0x6b25b57 (0x71c4c5e39b57 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #5: <unknown function> + 0x6a87113 (0x71c4c5d9b113 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #6: c10d::TCPStore::TCPStore(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, c10d::TCPStoreOptions const&) + 0x41d (0x71c4c5da1ead in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #7: <unknown function> + 0xecbf51 (0x71c4d5872f51 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so)
frame #8: <unknown function> + 0x411060 (0x71c4d4db8060 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so)
frame #9: /usr/bin/python3() [0x581e9f]
frame #10: _PyObject_MakeTpCall + 0x13e (0x54903e in /usr/bin/python3)
...snip...
frame #29: <unknown function> + 0x2a1ca (0x71c59dc851ca in /lib/x86_64-linux-gnu/libc.so.6)
frame #30: __libc_start_main + 0x8b (0x71c59dc8528b in /lib/x86_64-linux-gnu/libc.so.6)
frame #31: _start + 0x25 (0x6576c5 in /usr/bin/python3)

[E406 17:56:55.628826474 socket.cpp:1028] [c10d] The client socket has timed out after 60000ms while trying to connect to (10.100.10.1, 29500).
[E406 17:56:55.628974143 TCPStore.cpp:328] [c10d] TCP client failed to connect/validate to host 10.100.10.1:29500 - timed out (try=1, timeout=60000ms): The client socket has timed out after 60000ms while trying to connect to (10.100.10.1, 29500).
Exception raised from throwTimeoutError at /pytorch/torch/csrc/distributed/c10d/socket.cpp:1030 (most recent call first):
frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x71c56a97205d in /usr/local/lib/python3.12/dist-packages/torch/lib/libc10.so)
frame #1: <unknown function> + 0x16daa1c (0x71c4c09eea1c in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #2: <unknown function> + 0x6b25497 (0x71c4c5e39497 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #3: <unknown function> + 0x6b256cf (0x71c4c5e396cf in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #4: <unknown function> + 0x6b25b57 (0x71c4c5e39b57 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #5: <unknown function> + 0x6a87113 (0x71c4c5d9b113 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #6: c10d::TCPStore::TCPStore(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, c10d::TCPStoreOptions const&) + 0x41d (0x71c4c5da1ead in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #7: <unknown function> + 0xecbf51 (0x71c4d5872f51 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so)
frame #8: <unknown function> + 0x411060 (0x71c4d4db8060 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so)
frame #9: /usr/bin/python3() [0x581e9f]
frame #10: _PyObject_MakeTpCall + 0x13e (0x54903e in /usr/bin/python3)
...snip...
frame #30: __libc_start_main + 0x8b (0x71c59dc8528b in /lib/x86_64-linux-gnu/libc.so.6)
frame #31: _start + 0x25 (0x6576c5 in /usr/bin/python3)

Traceback (most recent call last):
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py", line 156, in _create_tcp_store
    store = TCPStore(
            ^^^^^^^^^
torch.distributed.DistNetworkError: The client socket has timed out after 60000ms while trying to connect to (10.100.10.1, 29500).

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
  File "/usr/local/bin/torchrun", line 6, in <module>
    sys.exit(main())
             ^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 362, in wrapper
    return f(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/run.py", line 990, in main
    run(args)
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/run.py", line 981, in run
    elastic_launch(
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/launcher/api.py", line 170, in __call__
    return launch_agent(self._config, self._entrypoint, list(args))
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/launcher/api.py", line 284, in launch_agent
    rdzv_handler=rdzv_registry.get_rendezvous_handler(rdzv_parameters),
                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/registry.py", line 96, in get_rendezvous_handler
    return handler_registry.create_handler(params)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/api.py", line 377, in create_handler
    handler = creator(params)
              ^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/registry.py", line 46, in _create_c10d_handler
    backend, store = create_backend(params)
                     ^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py", line 254, in create_backend
    store = _create_tcp_store(params)
            ^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py", line 180, in _create_tcp_store
    raise RendezvousConnectionError(
torch.distributed.elastic.rendezvous.api.RendezvousConnectionError: The connection to the C10d store has failed. See inner exception for details.
```

It failed with timeout! It seems that the process cannot connect to the rendezvous point 10.100.10.1:29500.

The log says `throwTimeoutError` was raised at [this line](https://github.com/pytorch/pytorch/blob/v2.11.0/torch/csrc/distributed/c10d/socket.cpp#L1030).  Is it time to file an issue there? No, it’s not what this blog does. It’s time to attach debugger!

As you may know, `torchrun` is just a python script to kick `torch.distributed.run`.

```log
root@2f3d44d8b0ea:/app# which torchrun
/usr/local/bin/torchrun
root@2f3d44d8b0ea:/app# cat /usr/local/bin/torchrun
#!/usr/bin/python3
import sys
from torch.distributed.run import main
if __name__ == '__main__':
    sys.argv[0] = sys.argv[0].removesuffix('.exe')
    sys.exit(main())
```

To debug it, you launch python with gdb, run the script, and break it when it’s stuck before the timeout exception is thrown.

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 gdb -q /usr/bin/python3
Reading symbols from /usr/bin/python3...
(No debugging symbols found in /usr/bin/python3)
(gdb) set pagination off
(gdb) r /usr/local/bin/torchrun --nnodes=1 --nproc_per_node=1 --node_rank=0 --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500 train.py
Starting program: /usr/bin/python3 /usr/local/bin/torchrun --nnodes=1 --nproc_per_node=1 --node_rank=0 --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500 train.py
warning: Error disabling address space randomization: Operation not permitted
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".
[New Thread 0x7733f25ff6c0 (LWP 1528)]
...snip...
[New Thread 0x7733e55e56c0 (LWP 1554)]
I0406 18:11:48.680000 1525 torch/distributed/run.py:735] Using nproc_per_node=1.
[New Thread 0x7733dfe2a6c0 (LWP 1555)]
I0406 18:11:48.788000 1525 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'timeout': 900}
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_iju9cg3h
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 18:11:48.789000 1525 torch/distributed/launcher/api.py:224]
^C
Thread 1 "pt_elastic" received signal SIGINT, Interrupt.
0x0000773541c2dadf in __GI___clock_nanosleep (clock_id=clock_id@entry=0, flags=flags@entry=0, req=0x7ffc61d9df30, rem=0x0)
    at ../sysdeps/unix/sysv/linux/clock_nanosleep.c:78
warning: 78     ../sysdeps/unix/sysv/linux/clock_nanosleep.c: No such file or directory
(gdb) bt
#0  0x0000773541c2dadf in __GI___clock_nanosleep (clock_id=clock_id@entry=0, flags=flags@entry=0, req=0x7ffc61d9df30, rem=0x0)
    at ../sysdeps/unix/sysv/linux/clock_nanosleep.c:78
#1  0x0000773541c3aa27 in __GI___nanosleep (req=<optimized out>, rem=<optimized out>) at ../sysdeps/unix/sysv/linux/nanosleep.c:25
#2  0x0000773469c3849a in c10d::detail::(anonymous namespace)::SocketConnectOp::tryConnect(int) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#3  0x0000773469c396cf in c10d::detail::(anonymous namespace)::SocketConnectOp::run() [clone .constprop.0] ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#4  0x0000773469c39b57 in c10d::detail::Socket::connect(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, c10d::detail::SocketOptions const&) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#5  0x0000773469b9b113 in c10d::detail::TCPClient::connect(c10d::detail::SocketAddress const&, c10d::TCPStoreOptions const&, std::shared_ptr<c10d::Backoff>)
    () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#6  0x0000773469ba1ead in c10d::TCPStore::TCPStore(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, c10d::TCPStoreOptions const&) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#7  0x0000773479672f51 in pybind11::cpp_function::initialize<pybind11::detail::initimpl::factory<torch::distributed::c10d::(anonymous namespace)::c10d_init(_object*, _object*)::{lambda(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool)#1}, pybind11::detail::void_type (*)(), c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > (std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool), pybind11::detail::void_type ()>::execute<pybind11::class_<c10d::TCPStore, c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > >, pybind11::arg, pybind11::arg, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, char [24]>(pybind11::class_<c10d::TCPStore, c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > >&, pybind11::arg const&, pybind11::arg const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, char const (&) [24]) &&::{lambda(pybind11::detail::value_and_holder&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool)#1}, void, pybind11::detail::value_and_holder&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool, pybind11::name, pybind11::is_method, pybind11::sibling, pybind11::detail::is_new_style_constructor, pybind11::arg, pybind11::arg, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, char [24]>(pybind11::class_<c10d::TCPStore, c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > >&&, void (*)(pybind11::detail::value_and_holder&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool), pybind11::name const&, pybind11::is_method const&, pybind11::sibling const&, pybind11::detail::is_new_style_constructor const&, pybind11::arg const&, pybind11::arg const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, char const (&) [24])::{lambda(pybind11::detail::function_call&)#1}::_FUN(pybind11::detail::function_call&) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so
#8  0x0000773478bb8060 in pybind11::cpp_function::dispatcher(_object*, _object*, _object*) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so
...snip...
#28 0x00000000006bc3ed in Py_BytesMain ()
#29 0x0000773541b6b1ca in __libc_start_call_main (main=main@entry=0x518930, argc=argc@entry=9, argv=argv@entry=0x7ffc61d9fdf8)
    at ../sysdeps/nptl/libc_start_call_main.h:58
#30 0x0000773541b6b28b in __libc_start_main_impl (main=0x518930, argc=9, argv=0x7ffc61d9fdf8, init=<optimized out>, fini=<optimized out>,
    rtld_fini=<optimized out>, stack_end=0x7ffc61d9fde8) at ../csu/libc-start.c:360
#31 0x00000000006576c5 in _start ()
(gdb)
```

It’s running [`SocketConnectOp::tryConnect`](https://github.com/pytorch/pytorch/blob/v2.11.0/torch/csrc/distributed/c10d/socket.cpp#L824) , which calls the standard `connect` function via `tryConnectCore` , that is expected to fail.  Let’s double check. You can just set a breakpoint there.

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 gdb -q /usr/bin/python3
Reading symbols from /usr/bin/python3...
(No debugging symbols found in /usr/bin/python3)
(gdb) set pagination off
(gdb) b SocketConnectOp::tryConnect
Function "SocketConnectOp::tryConnect" not defined.
Make breakpoint pending on future shared library load? (y or [n]) y
Breakpoint 1 (SocketConnectOp::tryConnect) pending.
(gdb) r /usr/local/bin/torchrun --nnodes=1 --nproc_per_node=1 --node_rank=0 --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500 train.py
Starting program: /usr/bin/python3 /usr/local/bin/torchrun --nnodes=1 --nproc_per_node=1 --node_rank=0 --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500 train.py
warning: Error disabling address space randomization: Operation not permitted
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".
[New Thread 0x73fda9dff6c0 (LWP 1786)]
...
[New Thread 0x73fd9cde56c0 (LWP 1812)]
I0406 19:03:29.615000 1783 torch/distributed/run.py:735] Using nproc_per_node=1.
[New Thread 0x73fd977156c0 (LWP 1813)]
I0406 19:03:29.701000 1783 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'timeout': 900}
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_wofb5igg
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 19:03:29.702000 1783 torch/distributed/launcher/api.py:224]

Thread 1 "pt_elastic" hit Breakpoint 1, 0x000073fe214380b0 in c10d::detail::(anonymous namespace)::SocketConnectOp::tryConnect(int) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
(gdb) b connect
Breakpoint 2 at 0x73fe2139b070 (17 locations)
(gdb) c
Continuing.

Thread 1 "pt_elastic" hit Breakpoint 2.17, __libc_connect (fd=24, addr=addr@entry=..., len=len@entry=16) at ../sysdeps/unix/sysv/linux/connect.c:24
warning: 24     ../sysdeps/unix/sysv/linux/connect.c: No such file or directory
(gdb) bt
#0  __libc_connect (fd=24, addr=addr@entry=..., len=len@entry=16) at ../sysdeps/unix/sysv/linux/connect.c:24
#1  0x000073fef94f7d04 in reopen (statp=statp@entry=0x73fef95b8680 <_res>, terrno=terrno@entry=0x7ffd0220c058, ns=ns@entry=0) at ./resolv/res_send.c:856
#2  0x000073fef94f88ee in send_dg (ansp2_malloced=<optimized out>, resplen2=<optimized out>, anssizp2=<optimized out>, ansp2=<optimized out>,
    anscp=<optimized out>, gotsomewhere=<synthetic pointer>, v_circuit=<synthetic pointer>, ns=<optimized out>, terrno=0x7ffd0220c058,
    anssizp=0x7ffd0220c190, ansp=0x7ffd0220c048, buflen2=<optimized out>, buf2=<optimized out>, buflen=<optimized out>, buf=<optimized out>,
    statp=<optimized out>) at ./resolv/res_send.c:957
#3  __GI___res_context_send (ctx=ctx@entry=0x187d3590, buf=buf@entry=0x7ffd0220c240 "\267\252\001", buflen=<optimized out>, buf2=buf2@entry=0x0,
    buflen2=buflen2@entry=0, ans=<optimized out>, ans@entry=0x7ffd0220ca80 "", anssiz=<optimized out>, ansp=<optimized out>, ansp2=<optimized out>,
    nansp2=<optimized out>, resplen2=<optimized out>, ansp2_malloced=<optimized out>) at ./resolv/res_send.c:373
#4  0x000073fef94f6217 in __GI___res_context_query (ctx=ctx@entry=0x187d3590, name=name@entry=0x7ffd0220ce80 "1.10.100.10.in-addr.arpa", class=class@entry=1,
    type=type@entry=12, answer=answer@entry=0x7ffd0220ca80 "", anslen=anslen@entry=1024, answerp=0x7ffd0220c718, answerp2=0x0, nanswerp2=0x0, resplen2=0x0,
    answerp2_malloced=0x0) at ./resolv/res_query.c:218
#5  0x000073fef94ef564 in __GI__nss_dns_gethostbyaddr2_r (addr=<optimized out>, len=<optimized out>, af=<optimized out>, result=0x7ffd0220d460,
    buffer=<optimized out>, buflen=<optimized out>, errnop=0x73fef93ac278, h_errnop=0x7ffd0220d444, ttlp=0x0) at nss_dns/dns-host.c:576
#6  0x000073fef94efa59 in __GI__nss_dns_gethostbyaddr_r (addr=<optimized out>, len=<optimized out>, af=<optimized out>, result=<optimized out>,
    buffer=<optimized out>, buflen=<optimized out>, errnop=0x73fef93ac278, h_errnop=0x7ffd0220d444) at nss_dns/dns-host.c:630
#7  0x000073fef9509f2c in __gethostbyaddr_r (addr=addr@entry=0x187b0d58, len=len@entry=16, type=type@entry=10, resbuf=resbuf@entry=0x7ffd0220d460,
    buffer=<optimized out>, buflen=<optimized out>, result=<optimized out>, h_errnop=<optimized out>) at ../nss/getXXbyYY_r.c:273
#8  0x000073fef950ba86 in gni_host_inet_name (addrlen=<optimized out>, flags=<optimized out>, hostlen=1025, host=0x7ffd0220df60 "", sa=0x187b0d50,
    tmpbuf=0x7ffd0220d4a0) at ./nss/getnameinfo.c:243
#9  gni_host_inet (addrlen=<optimized out>, flags=<optimized out>, hostlen=1025, host=0x7ffd0220df60 "", sa=0x187b0d50, tmpbuf=0x7ffd0220d4a0)
    at ./nss/getnameinfo.c:382
#10 gni_host (addrlen=<optimized out>, flags=<optimized out>, hostlen=<optimized out>, host=<optimized out>, sa=<optimized out>, tmpbuf=<optimized out>)
    at ./nss/getnameinfo.c:424
#11 __GI_getnameinfo (sa=0x187b0d50, addrlen=<optimized out>, host=0x7ffd0220df60 "", hostlen=1025, serv=0x7ffd0220dd70 "", servlen=32, flags=<optimized out>)
    at ./nss/getnameinfo.c:538
#12 0x000073fe214368bb in c10d::detail::formatSockAddr[abi:cxx11](sockaddr const*, unsigned int) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#13 0x000073fe2143c03d in void fmt::v12::detail::value<fmt::v12::context>::format_custom<addrinfo>(void*, fmt::v12::parse_context<char>&, fmt::v12::context&)
    () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#14 0x000073fe1c2fba18 in fmt::v12::vformat[abi:cxx11](fmt::v12::basic_string_view<char>, fmt::v12::basic_format_args<fmt::v12::context>) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#15 0x000073fe21436748 in c10d::detail::SocketImpl::SocketImpl(int, addrinfo const&) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#16 0x000073fe21438990 in c10d::detail::(anonymous namespace)::SocketConnectOp::tryConnect(int) ()
...
(gdb) c
Continuing.

Thread 1 "pt_elastic" hit Breakpoint 2.17, __libc_connect (fd=23, addr=..., len=28) at ../sysdeps/unix/sysv/linux/connect.c:24
24      in ../sysdeps/unix/sysv/linux/connect.c
(gdb) bt
#0  __libc_connect (fd=23, addr=..., len=28) at ../sysdeps/unix/sysv/linux/connect.c:24
#1  0x000073fe214389e3 in c10d::detail::(anonymous namespace)::SocketConnectOp::tryConnect(int) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#2  0x000073fe214396cf in c10d::detail::(anonymous namespace)::SocketConnectOp::run() [clone .constprop.0] ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#3  0x000073fe21439b57 in c10d::detail::Socket::connect(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, c10d::detail::SocketOptions const&) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#4  0x000073fe2139b113 in c10d::detail::TCPClient::connect(c10d::detail::SocketAddress const&, c10d::TCPStoreOptions const&, std::shared_ptr<c10d::Backoff>)
    () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#5  0x000073fe213a1ead in c10d::TCPStore::TCPStore(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, c10d::TCPStoreOptions const&) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#6  0x000073fe30e72f51 in pybind11::cpp_function::initialize<pybind11::detail::initimpl::factory<torch::distributed::c10d::(anonymous namespace)::c10d_init(_object*, _object*)::{lambda(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool)#1}, pybind11::detail::void_type (*)(), c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > (std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool), pybind11::detail::void_type ()>::execute<pybind11::class_<c10d::TCPStore, c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > >, pybind11::arg, pybind11::arg, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, char [24]>(pybind11::class_<c10d::TCPStore, c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > >&, pybind11::arg const&, pybind11::arg const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, char const (&) [24]) &&::{lambda(pybind11::detail::value_and_holder&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool)#1}, void, pybind11::detail::value_and_holder&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool, pybind11::name, pybind11::is_method, pybind11::sibling, pybind11::detail::is_new_style_constructor, pybind11::arg, pybind11::arg, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, char [24]>(pybind11::class_<c10d::TCPStore, c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > >&&, void (*)(pybind11::detail::value_and_holder&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool), pybind11::name const&, pybind11::is_method const&, pybind11::sibling const&, pybind11::detail::is_new_style_constructor const&, pybind11::arg const&, pybind11::arg const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, char const (&) [24])::{lambda(pybind11::detail::function_call&)#1}::_FUN(pybind11::detail::function_call&) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so
#7  0x000073fe303b8060 in pybind11::cpp_function::dispatcher(_object*, _object*, _object*) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so
...
(gdb) b *0x000073fe214389e3
Breakpoint 3 at 0x73fe214389e3
(gdb) c
Continuing.

Thread 1 "pt_elastic" hit Breakpoint 3, 0x000073fe214389e3 in c10d::detail::(anonymous namespace)::SocketConnectOp::tryConnect(int) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
(gdb) p $eax
$1 = -1
(gdb) p (int)errno
$2 = 115

```

It hit twice. The first one is from `getnameinfo`, which looks unrelated and we skip. The second one was called from `SocketConnectOp::tryConnect` and failed with `EINPROGRESS` (= 115).  This is the one we're interested in. And we’re interested in the parameters of `connect`. Here’s assembly of where we are, immediately after the call to `connect`.

```log
(gdb) disas 0x000073fe214380b0
...
   0x000073fe214389ca <+2330>:  mov    -0x5a8(%rbp),%rax
   0x000073fe214389d1 <+2337>:  mov    0x10(%rax),%edx
   0x000073fe214389d4 <+2340>:  mov    0x18(%rax),%rsi
   0x000073fe214389d8 <+2344>:  mov    0x50(%rbx),%rax
   0x000073fe214389dc <+2348>:  mov    (%rax),%edi
   0x000073fe214389de <+2350>:  call   0x73fe1b9afb00 <connect@plt>
=> 0x000073fe214389e3 <+2355>:  test   %eax,%eax
```

What we want is the 2nd parameter `const struct sockaddr *addr`, which is passed via `$rsi`  in [System V ABI](https://wiki.osdev.org/System_V_ABI).  Since we already lost `$rdi`, we need restore it from the stack.

```log
(gdb) x/1g $rbp-0x5a8
0x7ffd0220e7c8: 0x00000000187b0d20
(gdb) x/1g 0x00000000187b0d20+0x18
0x187b0d38:     0x00000000187b0d50
(gdb) x/28xb 0x00000000187b0d50
0x187b0d50:     0x0a    0x00    0x73    0x3c    0x00    0x00    0x00    0x00
0x187b0d58:     0x00    0x00    0x00    0x00    0x00    0x00    0x00    0x00
0x187b0d60:     0x00    0x00    0xff    0xff    0x0a    0x64    0x0a    0x01
0x187b0d68:     0x00    0x00    0x00    0x00
```

How to read this?  The address family is `0x0a 0x00` , which means `AF_INET6`, and we can see the address is `0xff 0xff 0x0a 0x64 0x0a 0x01`, which is an IPv4-mapped IPv6 address `::ffff:10.100.10.1`. The port is `0x73 0x3c` , which is 0x733c=29500.  This means the script simply tries to connect to the rendezvous point we specified, 10.100.10.1:29500, but it failed.  This means somebody should be listening on the endpoint.  Let's find out.

```
root@2f3d44d8b0ea:/app# ss -paln | grep 29500
root@2f3d44d8b0ea:/app#
```

Okay, nobody is listening on the endpoint, that's why `connect` failed.

The next thing to do is to see the positive behavior. We know this works with 127.0.0.1. Let’s see if the endpoint is listened on in that case.

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 gdb -q /usr/bin/python3
Reading symbols from /usr/bin/python3...
(No debugging symbols found in /usr/bin/python3)
(gdb) set pagination off
(gdb) b SocketConnectOp::tryConnect
Function "SocketConnectOp::tryConnect" not defined.
Make breakpoint pending on future shared library load? (y or [n]) y
Breakpoint 1 (SocketConnectOp::tryConnect) pending.
(gdb) r /usr/local/bin/torchrun --nnodes=1 --nproc_per_node=1 --node_rank=0 --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=127.0.0.1:29500 train.py
Starting program: /usr/bin/python3 /usr/local/bin/torchrun --nnodes=1 --nproc_per_node=1 --node_rank=0 --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=127.0.0.1:29500 train.py
warning: Error disabling address space randomization: Operation not permitted
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".
[New Thread 0x7b4c309ff6c0 (LWP 1830)]
...
[New Thread 0x7b4c239e56c0 (LWP 1856)]
I0406 19:22:02.314000 1827 torch/distributed/run.py:735] Using nproc_per_node=1.
[New Thread 0x7b4c1e32a6c0 (LWP 1857)]
I0406 19:22:02.402000 1827 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 127.0.0.1:29500
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'timeout': 900}
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_2oqcd4j3
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 19:22:02.403000 1827 torch/distributed/launcher/api.py:224]
[New Thread 0x7b4c10dff6c0 (LWP 1858)]

Thread 1 "pt_elastic" hit Breakpoint 1, 0x00007b4ca80380b0 in c10d::detail::(anonymous namespace)::SocketConnectOp::tryConnect(int) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
(gdb)
```

And see this! The endpoint is there.

```
root@2f3d44d8b0ea:/app# ss -paln | grep 29500
tcp   LISTEN 0      4096                              *:29500                *:*    users:(("pt_elastic",pid=1827,fd=30))
```

So the problem is the script doesn’t start listening on the endpoint if the address is 10.100.10.1, while it does on 127.0.0.1.

Do you know which function to set a breakpoint on? Probably `listen` or `bind`?  Well, in this case, I did some homework for you already and it turned out `bind` was the one. So let’s do it.

```log
(gdb) del
Delete all breakpoints, watchpoints, tracepoints, and catchpoints? (y or n) y
(gdb) b bind
Breakpoint 2 at 0x7b4caa493780 (9 locations)
(gdb) r /usr/local/bin/torchrun --nnodes=1 --nproc_per_node=1 --node_rank=0 --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=127.0.0.1:29500 train.py
The program being debugged has been started already.
Start it from the beginning? (y or n) y
Starting program: /usr/bin/python3 /usr/local/bin/torchrun --nnodes=1 --nproc_per_node=1 --node_rank=0 --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=127.0.0.1:29500 train.py
warning: Error disabling address space randomization: Operation not permitted
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".
[New Thread 0x70b1e8fff6c0 (LWP 1863)]
...
[New Thread 0x70b1dbfe56c0 (LWP 1889)]

Thread 1 "python3" hit Breakpoint 2.8, 0x000070b26fdda030 in torch::jit::slot_dict_impl<torch::jit::detail::ParameterPolicy>::bind(pybind11::module_ const&, char const*) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so
(gdb) c
Continuing.

Thread 1 "python3" hit Breakpoint 2.7, 0x000070b26fdd9820 in torch::jit::slot_dict_impl<torch::jit::detail::BufferPolicy>::bind(pybind11::module_ const&, char const*) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so
(gdb) c
Continuing.

Thread 1 "python3" hit Breakpoint 2.6, 0x000070b26fdd9010 in torch::jit::slot_dict_impl<torch::jit::detail::ModulePolicy>::bind(pybind11::module_ const&, char const*) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so
(gdb) c
Continuing.
I0406 19:27:40.820000 1862 torch/distributed/run.py:735] Using nproc_per_node=1.
[New Thread 0x70b1d682a6c0 (LWP 1890)]

Thread 1 "pt_elastic" hit Breakpoint 2.9, __GI_bind () at ../sysdeps/unix/syscall-template.S:120
warning: 120    ../sysdeps/unix/syscall-template.S: No such file or directory
(gdb) bt
#0  __GI_bind () at ../sysdeps/unix/syscall-template.S:120
#1  0x000070b300eca4eb in ?? () from /lib/x86_64-linux-gnu/libcuda.so.1
#2  0x000070b300f18602 in ?? () from /lib/x86_64-linux-gnu/libcuda.so.1
#3  0x000070b337a34a76 in ?? () from /usr/local/lib/python3.12/dist-packages/torch/lib/../../nvidia/cuda_runtime/lib/libcudart.so.12
#4  0x000070b337a39638 in ?? () from /usr/local/lib/python3.12/dist-packages/torch/lib/../../nvidia/cuda_runtime/lib/libcudart.so.12
#5  0x000070b338605ed3 in __pthread_once_slow (once_control=0x70b337cb2220, init_routine=0x70b337a395f0) at ./nptl/pthread_once.c:116
#6  0x000070b337a88ed9 in ?? () from /usr/local/lib/python3.12/dist-packages/torch/lib/../../nvidia/cuda_runtime/lib/libcudart.so.12
#7  0x000070b337a356ff in ?? () from /usr/local/lib/python3.12/dist-packages/torch/lib/../../nvidia/cuda_runtime/lib/libcudart.so.12
#8  0x000070b337a4d99a in cudaGetDeviceCount () from /usr/local/lib/python3.12/dist-packages/torch/lib/../../nvidia/cuda_runtime/lib/libcudart.so.12
#9  0x000070b3379ad842 in c10::cuda::device_count() () from /usr/local/lib/python3.12/dist-packages/torch/lib/libc10_cuda.so
#10 0x000070b26ffed402 in THCPModule_getDeviceCount_wrap(_object*, _object*) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so
#11 0x0000000000581a8a in ?? ()
...
#25 0x00000000006bc3ed in Py_BytesMain ()
#26 0x000070b33858e1ca in __libc_start_call_main (main=main@entry=0x518930, argc=argc@entry=9, argv=argv@entry=0x7ffd9cdf2e68)
    at ../sysdeps/nptl/libc_start_call_main.h:58
#27 0x000070b33858e28b in __libc_start_main_impl (main=0x518930, argc=9, argv=0x7ffd9cdf2e68, init=<optimized out>, fini=<optimized out>,
    rtld_fini=<optimized out>, stack_end=0x7ffd9cdf2e58) at ../csu/libc-start.c:360
#28 0x00000000006576c5 in _start ()
(gdb) c
Continuing.
I0406 19:27:50.442000 1862 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 127.0.0.1:29500
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'timeout': 900}
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_lqlqv2qz
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 19:27:50.443000 1862 torch/distributed/launcher/api.py:224]

Thread 1 "pt_elastic" hit Breakpoint 2.9, __GI_bind () at ../sysdeps/unix/syscall-template.S:120
120     in ../sysdeps/unix/syscall-template.S

(gdb) bt
#0  __GI_bind () at ../sysdeps/unix/syscall-template.S:120
#1  0x000070b26b907b2d in uv.tcp_bind () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#2  0x000070b2605c0918 in c10d::detail::UvTcpServer::makeWithPort(uv_loop_s*, unsigned short, bool) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#3  0x000070b2605b78fe in c10d::detail::LibUVStoreDaemon::init(c10d::TCPStoreOptions const&) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#4  0x000070b2605b793f in c10d::detail::create_libuv_tcpstore_backend(c10d::TCPStoreOptions const&) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#5  0x000070b26059dc85 in c10d::detail::TCPServer::start(c10d::TCPStoreOptions const&) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#6  0x000070b2605a1bf2 in c10d::TCPStore::TCPStore(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, c10d::TCPStoreOptions const&) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#7  0x000070b270072f51 in pybind11::cpp_function::initialize<pybind11::detail::initimpl::factory<torch::distributed::c10d::(anonymous namespace)::c10d_init(_object*, _object*)::{lambda(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool)#1}, pybind11::detail::void_type (*)(), c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > (std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool), pybind11::detail::void_type ()>::execute<pybind11::class_<c10d::TCPStore, c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > >, pybind11::arg, pybind11::arg, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, char [24]>(pybind11::class_<c10d::TCPStore, c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > >&, pybind11::arg const&, pybind11::arg const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, char const (&) [24]) &&::{lambda(pybind11::detail::value_and_holder&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool)#1}, void, pybind11::detail::value_and_holder&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool, pybind11::name, pybind11::is_method, pybind11::sibling, pybind11::detail::is_new_style_constructor, pybind11::arg, pybind11::arg, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, char [24]>(pybind11::class_<c10d::TCPStore, c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > >&&, void (*)(pybind11::detail::value_and_holder&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool), pybind11::name const&, pybind11::is_method const&, pybind11::sibling const&, pybind11::detail::is_new_style_constructor const&, pybind11::arg const&, pybind11::arg const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, char const (&) [24])::{lambda(pybind11::detail::function_call&)#1}::_FUN(pybind11::detail::function_call&) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so
#8  0x000070b26f5b8060 in pybind11::cpp_function::dispatcher(_object*, _object*, _object*) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so
...
#28 0x00000000006bc3ed in Py_BytesMain ()
#29 0x000070b33858e1ca in __libc_start_call_main (main=main@entry=0x518930, argc=argc@entry=9, argv=argv@entry=0x7ffd9cdf2e68)
    at ../sysdeps/nptl/libc_start_call_main.h:58
#30 0x000070b33858e28b in __libc_start_main_impl (main=0x518930, argc=9, argv=0x7ffd9cdf2e68, init=<optimized out>, fini=<optimized out>,
    rtld_fini=<optimized out>, stack_end=0x7ffd9cdf2e58) at ../csu/libc-start.c:360
#31 0x00000000006576c5 in _start ()
```

Okay, we got it. In this positive scenario, we start listening on the endpoint through `TCPServer::start`.

Now, what happens if the address is 10.100.10.1? If you look at the debugger output earlier carefully, we called `connect` inside `TCPStore::TCPStore` through `TCPClient::connect`. So we know `TCPStore::TCPStore` is surely called. Let’s see if we call `TCPServer::start` or not.

```log
(gdb) del
Delete all breakpoints, watchpoints, tracepoints, and catchpoints? (y or n) y
(gdb) b TCPStore::TCPStore
Breakpoint 3 at 0x70b2605a1a90
(gdb) b TCPServer::start
Breakpoint 4 at 0x70b26059d990
(gdb) b TCPClient::connect
Breakpoint 5 at 0x70b26059b070
(gdb) r /usr/local/bin/torchrun --nnodes=1 --nproc_per_node=1 --node_rank=0 --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500 train.py
The program being debugged has been started already.
Start it from the beginning? (y or n) y
Starting program: /usr/bin/python3 /usr/local/bin/torchrun --nnodes=1 --nproc_per_node=1 --node_rank=0 --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500 train.py
warning: Error disabling address space randomization: Operation not permitted
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".
[New Thread 0x7b62f01ff6c0 (LWP 1900)]
...
[New Thread 0x7b62e31e56c0 (LWP 1926)]
I0406 19:36:25.975000 1899 torch/distributed/run.py:735] Using nproc_per_node=1.
[New Thread 0x7b62dda2a6c0 (LWP 1927)]
I0406 19:36:26.088000 1899 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'timeout': 900}
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_e8gckjjs
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 19:36:26.091000 1899 torch/distributed/launcher/api.py:224]

Thread 1 "pt_elastic" hit Breakpoint 3, 0x00007b63677a1a90 in c10d::TCPStore::TCPStore(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, c10d::TCPStoreOptions const&) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
(gdb) c
Continuing.

Thread 1 "pt_elastic" hit Breakpoint 5, 0x00007b636779b070 in c10d::detail::TCPClient::connect(c10d::detail::SocketAddress const&, c10d::TCPStoreOptions const&, std::shared_ptr<c10d::Backoff>) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
(gdb)
```

See? We reached `TCPClient::connect` without hitting `TCPServer::start`. This is the problem. Let’s go check the code and see how we call `TCPServer::start`.  Code is [here](https://github.com/pytorch/pytorch/blob/v2.11.0/torch/csrc/distributed/c10d/TCPStore.cpp#L269).  It’s behind the check `if (opts.isServer) {`.  Does this mean `isServer` was `false` in our case?  Let's confirm it on debugger.

```log
(gdb) del
Delete all breakpoints, watchpoints, tracepoints, and catchpoints? (y or n) y
(gdb) b TCPStore::TCPStore
Breakpoint 6 at 0x7b63677a1a90
(gdb) r /usr/local/bin/torchrun --nnodes=1 --nproc_per_node=1 --node_rank=0 --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500 train.py
The program being debugged has been started already.
Start it from the beginning? (y or n) y
Starting program: /usr/bin/python3 /usr/local/bin/torchrun --nnodes=1 --nproc_per_node=1 --node_rank=0 --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500 train.py
warning: Error disabling address space randomization: Operation not permitted
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".
[New Thread 0x71c1955ff6c0 (LWP 1929)]
...
[New Thread 0x71c1885e56c0 (LWP 1955)]
I0406 19:39:42.387000 1928 torch/distributed/run.py:735] Using nproc_per_node=1.
[New Thread 0x71c182f156c0 (LWP 1956)]
I0406 19:39:42.502000 1928 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'timeout': 900}
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_d435hy84
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 19:39:42.505000 1928 torch/distributed/launcher/api.py:224]

Thread 1 "pt_elastic" hit Breakpoint 6, 0x000071c20cba1a90 in c10d::TCPStore::TCPStore(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, c10d::TCPStoreOptions const&) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
(gdb) disas $rip
Dump of assembler code for function _ZN4c10d8TCPStoreC2ENSt7__cxx1112basic_stringIcSt11char_traitsIcESaIcEEERKNS_15TCPStoreOptionsE:
=> 0x000071c20cba1a90 <+0>:     push   %rbp
   0x000071c20cba1a91 <+1>:     pxor   %xmm0,%xmm0
...
   0x000071c20cba1bbf <+303>:   call   0x71c20cc32ba0 <_ZN4c10d6detail6Socket10initializeEv> <<<< Socket::initialize()
   0x000071c20cba1bc4 <+308>:   mov    -0x608(%rbp),%rsi
   0x000071c20cba1bcb <+315>:   movzwl (%rsi),%eax
   0x000071c20cba1bce <+318>:   cmpb   $0x0,0x2(%rsi) <<<< opts.isServer
   0x000071c20cba1bd2 <+322>:   mov    %ax,0x38(%rbx)
   0x000071c20cba1bd6 <+326>:   je     0x71c20cba1df1 <_ZN4c10d8TCPStoreC2ENSt7__cxx1112basic_stringIcSt11char_traitsIcESaIcEEERKNS_15TCPStoreOptionsE+865>
   0x000071c20cba1bdc <+332>:   lea    -0x240(%rbp),%rax
   0x000071c20cba1be3 <+339>:   mov    %rax,%rdi
   0x000071c20cba1be6 <+342>:   mov    %rax,-0x620(%rbp)
   0x000071c20cba1bed <+349>:   call   0x71c20cb9d990 <_ZN4c10d6detail9TCPServer5startERKNS_15TCPStoreOptionsE> <<<< TCPServer::start
   0x000071c20cba1bf2 <+354>:   mov    0x48(%rbx),%rdi
   0x000071c20cba1bf6 <+358>:   movdqa -0x240(%rbp),%xmm3
   0x000071c20cba1bfe <+366>:   movups %xmm3,0x40(%rbx)
   0x000071c20cba1c02 <+370>:   test   %rdi,%rdi
...
```

As I commented inline, `cmpb $0x0,0x2(%rsi)` is checking the flag.

```
(gdb) b *0x000071c20cba1bce
Breakpoint 7 at 0x71c20cba1bce
(gdb) c
Continuing.

Thread 1 "pt_elastic" hit Breakpoint 7, 0x000071c20cba1bce in c10d::TCPStore::TCPStore(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, c10d::TCPStoreOptions const&) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
(gdb) x/1b $rsi+2
0x7ffcc728a192: 0
```

It’s zero! This is why we skip listening on the rendezvous point!

And it’s from the parameter `opts`. Where does it come from?  Who instantiates `TCPStore` ?

```log
(gdb) bt
#0  0x000071c20cba1bce in c10d::TCPStore::TCPStore(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, c10d::TCPStoreOptions const&) () from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so
#1  0x000071c21c672f51 in pybind11::cpp_function::initialize<pybind11::detail::initimpl::factory<torch::distributed::c10d::(anonymous namespace)::c10d_init(_object*, _object*)::{lambda(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool)#1}, pybind11::detail::void_type (*)(), c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > (std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool), pybind11::detail::void_type ()>::execute<pybind11::class_<c10d::TCPStore, c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > >, pybind11::arg, pybind11::arg, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, char [24]>(pybind11::class_<c10d::TCPStore, c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > >&, pybind11::arg const&, pybind11::arg const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, char const (&) [24]) &&::{lambda(pybind11::detail::value_and_holder&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool)#1}, void, pybind11::detail::value_and_holder&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool, pybind11::name, pybind11::is_method, pybind11::sibling, pybind11::detail::is_new_style_constructor, pybind11::arg, pybind11::arg, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, pybind11::arg_v, char [24]>(pybind11::class_<c10d::TCPStore, c10::intrusive_ptr<c10d::TCPStore, c10::detail::intrusive_target_default_null_type<c10d::TCPStore> > >&&, void (*)(pybind11::detail::value_and_holder&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, unsigned short, std::optional<int>, bool, std::chrono::duration<long, std::ratio<1l, 1000l> >, bool, bool, std::optional<int>, bool), pybind11::name const&, pybind11::is_method const&, pybind11::sibling const&, pybind11::detail::is_new_style_constructor const&, pybind11::arg const&, pybind11::arg const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, pybind11::arg_v const&, char const (&) [24])::{lambda(pybind11::detail::function_call&)#1}::_FUN(pybind11::detail::function_call&) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so
#2  0x000071c21bbb8060 in pybind11::cpp_function::dispatcher(_object*, _object*, _object*) ()
   from /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so
#3  0x0000000000581e9f in ?? ()
```

The symbol of frame #1 is insanely long.  See `pybind11` namespace. It’s python binding, meaning it’s python script instantiating `TCPStore` in C++.  `opts.isServer`  is also from python.

### Debugging with PDB

How to debug Python? Do we need VSCode or some fancy IDE to debug it? That’s not what this blog does.  Do we debug python binding?  Well, that's too ambitious.

One of the great functionality Python provides, compared to other script languages, is Python has a built-in console debugger, [pdb](https://docs.python.org/3/library/pdb.html). You don’t need additional components to live debug python code.  Very handy.

To use pdb, you modify your script to break at the beginning. Since we’re running our script with `torchrun`, we make a little modification in `torchrun` itself, just adding one line `breakpoint()` before `main()`.

```python
root@2f3d44d8b0ea:/app# cat /usr/local/bin/torchrun
#!/usr/bin/python3
import sys
from torch.distributed.run import main
if __name__ == '__main__':
    breakpoint() # <---- add this line
    sys.argv[0] = sys.argv[0].removesuffix('.exe')
    sys.exit(main())
```

Where does it call C++? A formal approach would be to start from the log `Starting elastic_operator with launch configs:`, which is from [this line](https://github.com/pytorch/pytorch/blob/v2.11.0/torch/distributed/launcher/api.py#L225). In our case, however, let’s simply search the repo for the keyword `TCPStore` .  This function [`_create_tcp_store`](https://github.com/pytorch/pytorch/blob/v2.11.0/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py#L156) looks suspicious.  Since python is a script language, you can set a breakpoint accurately with a source line.

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 torchrun   --nnodes=1 --nproc_per_node=1 --node_rank=0   --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500   train.py
> /usr/local/bin/torchrun(6)<module>()
-> sys.argv[0] = sys.argv[0].removesuffix('.exe')
(Pdb) b /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py:154
Breakpoint 1 at /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py:154
(Pdb) c
I0406 20:29:58.666000 1992 torch/distributed/run.py:735] Using nproc_per_node=1.
I0406 20:29:58.778000 1992 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'timeout': 900}
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_fw4c6ft_
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 20:29:58.781000 1992 torch/distributed/launcher/api.py:224]
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py(154)_create_tcp_store()
-> for is_server in [is_host, False]:
(Pdb) bt
  /usr/local/bin/torchrun(7)<module>()
-> sys.exit(main())
  /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py(362)wrapper()
-> return f(*args, **kwargs)
  /usr/local/lib/python3.12/dist-packages/torch/distributed/run.py(990)main()
-> run(args)
  /usr/local/lib/python3.12/dist-packages/torch/distributed/run.py(981)run()
-> elastic_launch(
  /usr/local/lib/python3.12/dist-packages/torch/distributed/launcher/api.py(170)__call__()
-> return launch_agent(self._config, self._entrypoint, list(args))
  /usr/local/lib/python3.12/dist-packages/torch/distributed/launcher/api.py(284)launch_agent()
-> rdzv_handler=rdzv_registry.get_rendezvous_handler(rdzv_parameters),
  /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/registry.py(96)get_rendezvous_handler()
-> return handler_registry.create_handler(params)
  /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/api.py(377)create_handler()
-> handler = creator(params)
  /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/registry.py(46)_create_c10d_handler()
-> backend, store = create_backend(params)
  /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py(254)create_backend()
-> store = _create_tcp_store(params)
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py(154)_create_tcp_store()
-> for is_server in [is_host, False]:
(Pdb) p is_host
False
(Pdb) p cfg_is_host
None
(Pdb) p host
'10.100.10.1'
```

Okay, we got it. `_create_tcp_store` is instantiating `TCPStore` with `is_master=is_server`, which is `False`.

Where does this `is_host` come from? Since `cfg_is_host` is None , it must be from [`_matches_machine_hostname`](https://github.com/pytorch/pytorch/blob/v2.11.0/torch/distributed/elastic/rendezvous/utils.py#L117).  Let’s run this function line by line.

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 torchrun   --nnodes=1 --nproc_per_node=1 --node_rank=0   --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500   train.py
> /usr/local/bin/torchrun(6)<module>()
-> sys.argv[0] = sys.argv[0].removesuffix('.exe')
(Pdb) b /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/utils.py:146
Breakpoint 1 at /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/utils.py:146
(Pdb) c
I0406 20:35:35.452000 2021 torch/distributed/run.py:735] Using nproc_per_node=1.
I0406 20:35:35.549000 2021 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'timeout': 900}
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_dek76zg_
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 20:35:35.551000 2021 torch/distributed/launcher/api.py:224]
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/utils.py(146)_matches_machine_hostname()
-> if host == this_host:
(Pdb) p host
'10.100.10.1'
(Pdb) p this_host
'2f3d44d8b0ea'
(Pdb) n
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/utils.py(149)_matches_machine_hostname()
-> addr_list = socket.getaddrinfo(
(Pdb)
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/utils.py(150)_matches_machine_hostname()
-> this_host, None, proto=socket.IPPROTO_TCP, flags=socket.AI_CANONNAME
(Pdb)
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/utils.py(149)_matches_machine_hostname()
-> addr_list = socket.getaddrinfo(
(Pdb)
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/utils.py(152)_matches_machine_hostname()
-> for addr_info in addr_list:
(Pdb) p addr_list
[(<AddressFamily.AF_INET: 2>, <SocketKind.SOCK_STREAM: 1>, 6, '2f3d44d8b0ea', ('172.17.0.2', 0))]
(Pdb)
```

We clearly see the problem. First, we get the hostname with `gethostname()`, which is `2f3d44d8b0ea` (It matches the container’s ID). And we get the IP address associated with the hostname via `getaddrinfo`, which returns only `172.17.0.2`, the one mapped to the default bridge interface (eth0). Since it doesn’t match `10.100.10.1`, it thinks “I’m not the master. Somebody else should start listening on the rendezvous endpoint.”

Ideally `_matches_machine_hostname` should iterate all IP addressed assigned to check any of them matches the host. We may consider sending a PR to them, but for now, is there a good way to work around this behavior?

There is. See the beginning of `_create_tcp_store`. It’s overwritable through `cfg_is_host`, coming from `params`.

```python
    cfg_is_host = params.get_as_bool("is_host")
    # If the user has explicitly specified whether our process should host the
    # the store, respect it.
    if cfg_is_host is not None:
        is_host = cfg_is_host
    # Otherwise try to determine whether we are the host based on our hostname
    # and IP address.
    else:
        is_host = _matches_machine_hostname(host)
```

Where is this `RendezvousParameters` created? It’s way above `_matches_machine_hostname`. It’s in the function [`run`](https://github.com/pytorch/pytorch/blob/v2.11.0/torch/distributed/run.py#L980). `torchrun` takes a rarely-used parameter `--rdzv_conf` where we can specify key-value pairs. What we want to specify is `is_host`. This looks promising. Let’s try it.

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 torchrun \
  --nnodes=1 --nproc_per_node=1 --node_rank=0 \
  --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500 \
  --rdzv_conf=is_host=1 \
  train.py
I0406 20:52:48.947000 2060 torch/distributed/run.py:735] Using nproc_per_node=1.
I0406 20:52:49.029000 2060 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'is_host': 'True', 'timeout': 900}
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_ccaq0pji
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 20:52:49.030000 2060 torch/distributed/launcher/api.py:224]
I0406 20:52:49.055000 2060 torch/distributed/elastic/agent/server/api.py:898] [default] starting workers for entrypoint: python3
I0406 20:52:49.056000 2060 torch/distributed/elastic/agent/server/api.py:693] [default] Rendezvous'ing worker group
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/api.py:539] [default] Rendezvous complete for workers. Result:
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/api.py:539]   restart_count=0
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/api.py:539]   master_addr=2f3d44d8b0ea
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/api.py:539]   master_port=36587
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/api.py:539]   group_rank=0
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/api.py:539]   group_world_size=1
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/api.py:539]   local_ranks=[0]
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/api.py:539]   role_ranks=[0]
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/api.py:539]   global_ranks=[0]
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/api.py:539]   role_world_sizes=[1]
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/api.py:539]   global_world_sizes=[1]
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/api.py:539]   event_log_handler=null
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/api.py:539]
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/api.py:701] [default] Starting worker group
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/local_elastic_agent.py:299] use_agent_store: True
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/local_elastic_agent.py:195] Environment variable 'TORCHELASTIC_ENABLE_FILE_TIMER' not found. Do not start FileTimerServer.
I0406 20:52:49.283000 2060 torch/distributed/elastic/agent/server/local_elastic_agent.py:239] Environment variable 'TORCHELASTIC_HEALTH_CHECK_PORT' not found. Do not start health check.
2f3d44d8b0ea:2092:2092 [0] NCCL INFO ENV/Plugin: Could not find: libnccl-env.so
2f3d44d8b0ea:2092:2092 [0] NCCL INFO Bootstrap: Using eth0:172.17.0.2<0>
2f3d44d8b0ea:2092:2092 [0] NCCL INFO cudaDriverVersion 12080
2f3d44d8b0ea:2092:2092 [0] NCCL INFO NCCL version 2.28.9+cuda12.9
2f3d44d8b0ea:2092:2092 [0] NCCL INFO Comm config Blocking set to 1
2f3d44d8b0ea:2092:2092 [0] NCCL INFO NET/Plugin: Could not find: libnccl-net.so
2f3d44d8b0ea:2092:2092 [0] NCCL INFO Failed to open libibverbs.so[.1]
2f3d44d8b0ea:2092:2092 [0] NCCL INFO transport/net_ib.cc:852 -> 3
2f3d44d8b0ea:2092:2092 [0] NCCL INFO Failed to initialize NET plugin IB
2f3d44d8b0ea:2092:2092 [0] NCCL INFO NET/Socket : Using [0]eth0:172.17.0.2<0> [1]wg0:10.100.10.1<0>
2f3d44d8b0ea:2092:2092 [0] NCCL INFO Initialized NET plugin Socket
2f3d44d8b0ea:2092:2092 [0] NCCL INFO Assigned NET plugin Socket to comm
2f3d44d8b0ea:2092:2092 [0] NCCL INFO Using network Socket
2f3d44d8b0ea:2092:2092 [0] NCCL INFO ncclCommInitRankConfig comm 0x230fcf20 rank 0 nranks 1 cudaDev 0 nvmlDev 0 busId 70 commId 0xda9011adca28b7b9 - Init START
2f3d44d8b0ea:2092:2092 [0] NCCL INFO RAS client listening socket at ::1<28028>
2f3d44d8b0ea:2092:2092 [0] NCCL INFO Bootstrap timings total 0.001490 (create 0.000058, send 0.000233, recv 0.000596, ring 0.000001, delay 0.000000)
2f3d44d8b0ea:2092:2092 [0] NCCL INFO NCCL_IGNORE_DISABLED_P2P set by environment to 1.
2f3d44d8b0ea:2092:2092 [0] NCCL INFO ncclTopoGetCpuAffinity: Affinity for GPU 0 is empty, ignoring. (GPU affinity =  ; CPU affinity = 0-27).
2f3d44d8b0ea:2092:2092 [0] NCCL INFO comm 0x230fcf20 rank 0 nRanks 1 nNodes 1 localRanks 1 localRank 0 MNNVL 0
2f3d44d8b0ea:2092:2092 [0] NCCL INFO Channel 00/64 : 0
...
2f3d44d8b0ea:2092:2092 [0] NCCL INFO Channel 63/64 : 0
2f3d44d8b0ea:2092:2092 [0] NCCL INFO Trees [0] -1/-1/-1->0->-1 [1] -1/-1/-1->0->-1 [2] -1/-1/-1->0->-1 [3] -1/-1/-1->0->-1 [4] -1/-1/-1->0->-1 [5] -1/-1/-1->0->-1 [6] -1/-1/-1->0->-1 [7] -1/-1/-1->0->-1 [8] -1/-1/-1->0->-1 [9] -1/-1/-1->0->-1 [10] -1/-1/-1->0->-1 [11] -1/-1/-1->0->-1 [12] -1/-1/-1->0->-1 [13] -1/-1/-1->0->-1 [14] -1/-1/-1->0->-1 [15] -1/-1/-1->0->-1 [16] -1/-1/-1->0->-1 [17] -1/-1/-1->0->-1 [18] -1/-1/-1->0->-1 [19] -1/-1/-1->0->-1 [20] -1/-1/-1->0->-1 [21] -1/-1/-1->0->-1 [22] -1/-1/-1->0->-1 [23] -1/-1/-1->0->-1 [24] -1/-1/-1->0->-1 [25] -1/-1/-1->0->-1 [26] -1/-1/-1->0->-1 [27] -1/-1/-1->0->-1 [28] -1/-1/-1->0->-1 [29] -1/-1/-1->0->-1 [30] -1/-1/-1->0->-1 [31] -1/-1/-1->0->-1 [32] -1/-1/-1->0->-1 [33] -1/-1/-1->0->-1 [34] -1/-1/-1->0->-1 [35] -1/-1/-1->0->-1 [36] -1/-1/-1->0->-1 [37] -1/-1/-1->0->-1 [38] -1/-1/-1->0->-1 [39] -1/-1/-1->0->-1 [40] -1/-1/-1->0->-1 [41] -1/-1/-1->0->-1 [42] -1/-1/-1->0->-1 [43] -1/-1/-1->0->-1 [44] -1/-1/-1->0->-1 [45] -1/-1/-1->0->-1 [46] -1/-1/-1->0->-1 [47
2f3d44d8b0ea:2092:2092 [0] NCCL INFO P2P Chunksize set to 524288
2f3d44d8b0ea:2092:2092 [0] NCCL INFO PROFILER/Plugin: Could not find: libnccl-profiler.so
2f3d44d8b0ea:2092:2092 [0] NCCL INFO Check P2P Type isAllDirectP2p 1 directMode 0 isAllCudaP2p 1
2f3d44d8b0ea:2092:2126 [0] NCCL INFO [Proxy Service] Device 0 CPU core 2
2f3d44d8b0ea:2092:2127 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 12
2f3d44d8b0ea:2092:2092 [0] NCCL INFO TUNER/Plugin: Could not find: libnccl-tuner.so
2f3d44d8b0ea:2092:2092 [0] NCCL INFO 64 coll channels, 64 collnet channels, 0 nvls channels, 64 p2p channels, 64 p2p channels per peer
2f3d44d8b0ea:2092:2092 [0] NCCL INFO CC Off, workFifoBytes 1048576
2f3d44d8b0ea:2092:2092 [0] NCCL INFO ncclCommInitRankConfig comm 0x230fcf20 rank 0 nranks 1 cudaDev 0 nvmlDev 0 busId 70 commId 0xda9011adca28b7b9 - Init COMPLETE
2f3d44d8b0ea:2092:2092 [0] NCCL INFO Init timings - ncclCommInitRankConfig: rank 0 nranks 1 total 0.17 (kernels 0.13, alloc 0.00, bootstrap 0.00, allgathers 0.00, topo 0.00, graphs 0.00, connections 0.02, rest 0.00)
[Rank 0] Starting training...
Step 0 | Loss: 1.3868
Step 10 | Loss: 1.5847
Step 20 | Loss: 1.5615
[Rank 0] Training complete.
2f3d44d8b0ea:2092:2092 [0] NCCL INFO comm 0x230fcf20 rank 0 nranks 1 cudaDev 0 busId 70 - Destroy COMPLETE
2f3d44d8b0ea:2092:2092 [0] NCCL INFO ENV/Plugin: Closing env plugin ncclEnvDefault
I0406 20:52:52.994000 2060 torch/distributed/elastic/https://github.com/pytorch/pytorch/blob/v2.11.0/torch/distributed/run.py#L1006agent/server/api.py:917] [default] worker group successfully finished. Waiting 300 seconds for other agents to finish.
I0406 20:52:52.994000 2060 torch/distributed/elastic/agent/server/api.py:970] Local worker group finished (WorkerState.SUCCEEDED). Waiting 300 seconds for other agents to finish
I0406 20:52:52.995000 2060 torch/distributed/elastic/agent/server/api.py:984] Done waiting for other agents. Elapsed: 0.00042891502380371094 seconds
root@2f3d44d8b0ea:/app#
```

It worked like a charm!

Now, it’s time to run a training on two nodes?

### Training on two nodes

This is the first node; master node, rank 0.

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 torchrun \
  --nnodes=2 --nproc_per_node=1 --node_rank=0 \
  --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500 \
  --rdzv_conf=is_host=1 \
  train.py
I0406 20:59:43.723000 2129 torch/distributed/run.py:735] Using nproc_per_node=1.
I0406 20:59:43.825000 2129 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   min_nodes                : 2
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   max_nodes                : 2
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'is_host': 'True', 'timeout': 900}
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_ji8a86n0
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 20:59:43.826000 2129 torch/distributed/launcher/api.py:224]
I0406 20:59:43.854000 2129 torch/distributed/elastic/agent/server/api.py:898] [default] starting workers for entrypoint: python3
I0406 20:59:43.855000 2129 torch/distributed/elastic/agent/server/api.py:693] [default] Rendezvous'ing worker group
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/api.py:539] [default] Rendezvous complete for workers. Result:
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/api.py:539]   restart_count=0
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/api.py:539]   master_addr=2f3d44d8b0ea
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/api.py:539]   master_port=44403
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/api.py:539]   group_rank=0
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/api.py:539]   group_world_size=2
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/api.py:539]   local_ranks=[0]
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/api.py:539]   role_ranks=[0]
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/api.py:539]   global_ranks=[0]
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/api.py:539]   role_world_sizes=[2]
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/api.py:539]   global_world_sizes=[2]
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/api.py:539]   event_log_handler=null
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/api.py:539]
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/api.py:701] [default] Starting worker group
I0406 20:59:51.030000 2129 torch/distributed/elastic/agent/server/local_elastic_agent.py:299] use_agent_store: True
I0406 20:59:51.031000 2129 torch/distributed/elastic/agent/server/local_elastic_agent.py:195] Environment variable 'TORCHELASTIC_ENABLE_FILE_TIMER' not found. Do not start FileTimerServer.
I0406 20:59:51.031000 2129 torch/distributed/elastic/agent/server/local_elastic_agent.py:239] Environment variable 'TORCHELASTIC_HEALTH_CHECK_PORT' not found. Do not start health check.
2f3d44d8b0ea:2161:2161 [0] NCCL INFO ENV/Plugin: Could not find: libnccl-env.so
2f3d44d8b0ea:2161:2161 [0] NCCL INFO Bootstrap: Using eth0:172.17.0.2<0>
2f3d44d8b0ea:2161:2161 [0] NCCL INFO cudaDriverVersion 12080
2f3d44d8b0ea:2161:2161 [0] NCCL INFO NCCL version 2.28.9+cuda12.9
2f3d44d8b0ea:2161:2161 [0] NCCL INFO Comm config Blocking set to 1
2f3d44d8b0ea:2161:2161 [0] NCCL INFO NET/Plugin: Could not find: libnccl-net.so
2f3d44d8b0ea:2161:2161 [0] NCCL INFO Failed to open libibverbs.so[.1]
2f3d44d8b0ea:2161:2161 [0] NCCL INFO transport/net_ib.cc:852 -> 3
2f3d44d8b0ea:2161:2161 [0] NCCL INFO Failed to initialize NET plugin IB
2f3d44d8b0ea:2161:2161 [0] NCCL INFO NET/Socket : Using [0]eth0:172.17.0.2<0> [1]wg0:10.100.10.1<0>
2f3d44d8b0ea:2161:2161 [0] NCCL INFO Initialized NET plugin Socket
2f3d44d8b0ea:2161:2161 [0] NCCL INFO Assigned NET plugin Socket to comm
2f3d44d8b0ea:2161:2161 [0] NCCL INFO Using network Socket
2f3d44d8b0ea:2161:2161 [0] NCCL INFO ncclCommInitRankConfig comm 0x3afce880 rank 0 nranks 2 cudaDev 0 nvmlDev 0 busId 70 commId 0x622b058bf1d66d35 - Init START
```

This is the second node; rank 1.  Ouch!  It failed!

```log
root@dd3bc9f2214d:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 torchrun \
  --nnodes=2 --nproc_per_node=1 --node_rank=1 \
  --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500 \
  train.py
I0406 20:59:49.909000 131 torch/distributed/run.py:735] Using nproc_per_node=1.
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   min_nodes                : 2
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   max_nodes                : 2
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'timeout': 900}
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_bxkvajbb
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 20:59:49.992000 131 torch/distributed/launcher/api.py:224]
I0406 20:59:50.023000 131 torch/distributed/elastic/agent/server/api.py:898] [default] starting workers for entrypoint: python3
I0406 20:59:50.023000 131 torch/distributed/elastic/agent/server/api.py:693] [default] Rendezvous'ing worker group
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/api.py:539] [default] Rendezvous complete for workers. Result:
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/api.py:539]   restart_count=0
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/api.py:539]   master_addr=2f3d44d8b0ea
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/api.py:539]   master_port=44403
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/api.py:539]   group_rank=1
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/api.py:539]   group_world_size=2
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/api.py:539]   local_ranks=[0]
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/api.py:539]   role_ranks=[1]
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/api.py:539]   global_ranks=[1]
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/api.py:539]   role_world_sizes=[2]
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/api.py:539]   global_world_sizes=[2]
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/api.py:539]   event_log_handler=null
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/api.py:539]
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/api.py:701] [default] Starting worker group
I0406 20:59:51.032000 131 torch/distributed/elastic/agent/server/local_elastic_agent.py:299] use_agent_store: True
I0406 20:59:51.033000 131 torch/distributed/elastic/agent/server/local_elastic_agent.py:195] Environment variable 'TORCHELASTIC_ENABLE_FILE_TIMER' not found. Do not start FileTimerServer.
I0406 20:59:51.033000 131 torch/distributed/elastic/agent/server/local_elastic_agent.py:239] Environment variable 'TORCHELASTIC_HEALTH_CHECK_PORT' not found. Do not start health check.
[W406 20:59:52.360496683 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 20:59:52.816683350 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 20:59:53.625858136 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 20:59:54.305711072 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 20:59:56.854604631 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 20:59:58.498181734 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:00:02.234917616 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:00:10.139460584 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:00:17.403585181 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:00:25.605543230 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:00:36.109063693 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:01:06.103270816 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:01:32.333383773 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:02:48.901471959 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:03:22.483614915 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:04:02.944138445 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:05:13.293170306 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:06:36.582248302 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:07:41.396091809 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:08:12.654681204 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:08:54.086316393 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[E406 21:08:54.086357761 socket.cpp:1028] [c10d] The client socket has timed out after 600000ms while trying to connect to (2f3d44d8b0ea, 44403).
[W406 21:08:54.086697964 TCPStore.cpp:340] [c10d] TCP client failed to connect/validate to host 2f3d44d8b0ea:44403 - retrying (try=0, timeout=600000ms, delay=74560ms): The client socket has timed out after 600000ms while trying to connect to (2f3d44d8b0ea, 44403).
Exception raised from throwTimeoutError at /pytorch/torch/csrc/distributed/c10d/socket.cpp:1030 (most recent call first):
frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x7cadc417205d in /usr/local/lib/python3.12/dist-packages/torch/lib/libc10.so)
frame #1: <unknown function> + 0x16daa1c (0x7cace79eea1c in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #2: <unknown function> + 0x6b25497 (0x7cacece39497 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #3: <unknown function> + 0x6b256cf (0x7cacece396cf in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #4: <unknown function> + 0x6b25b57 (0x7cacece39b57 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #5: <unknown function> + 0x6a87113 (0x7cacecd9b113 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #6: c10d::TCPStore::TCPStore(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, c10d::TCPStoreOptions const&) + 0x41d (0x7cacecda1ead in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #7: <unknown function> + 0xecbf51 (0x7cacfc872f51 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so)
frame #8: <unknown function> + 0x411060 (0x7cacfbdb8060 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so)
frame #9: /usr/bin/python3() [0x581e9f]
...
frame #32: _start + 0x25 (0x6576c5 in /usr/bin/python3)

[W406 21:10:08.720419645 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:10:48.462175738 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:11:42.498706496 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:12:12.849740252 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:12:43.260536804 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:13:41.678641445 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:14:14.137672492 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:15:41.308584356 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:16:48.530796444 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:17:19.359159001 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:18:45.552803370 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[W406 21:19:24.849525746 socket.cpp:764] [c10d] The IPv6 network addresses of (2f3d44d8b0ea, 44403) cannot be retrieved (gai error: -2 - Name or service not known).
[E406 21:19:24.849598964 socket.cpp:1028] [c10d] The client socket has timed out after 600000ms while trying to connect to (2f3d44d8b0ea, 44403).
[E406 21:19:24.849748526 TCPStore.cpp:328] [c10d] TCP client failed to connect/validate to host 2f3d44d8b0ea:44403 - timed out (try=1, timeout=600000ms): The client socket has timed out after 600000ms while trying to connect to (2f3d44d8b0ea, 44403).
Exception raised from throwTimeoutError at /pytorch/torch/csrc/distributed/c10d/socket.cpp:1030 (most recent call first):
frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0x9d (0x7cadc417205d in /usr/local/lib/python3.12/dist-packages/torch/lib/libc10.so)
frame #1: <unknown function> + 0x16daa1c (0x7cace79eea1c in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #2: <unknown function> + 0x6b25497 (0x7cacece39497 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #3: <unknown function> + 0x6b256cf (0x7cacece396cf in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #4: <unknown function> + 0x6b25b57 (0x7cacece39b57 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #5: <unknown function> + 0x6a87113 (0x7cacecd9b113 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #6: c10d::TCPStore::TCPStore(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, c10d::TCPStoreOptions const&) + 0x41d (0x7cacecda1ead in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_cpu.so)
frame #7: <unknown function> + 0xecbf51 (0x7cacfc872f51 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so)
frame #8: <unknown function> + 0x411060 (0x7cacfbdb8060 in /usr/local/lib/python3.12/dist-packages/torch/lib/libtorch_python.so)
frame #9: /usr/bin/python3() [0x581e9f]
...
frame #31: __libc_start_main + 0x8b (0x7cadc4e2128b in /lib/x86_64-linux-gnu/libc.so.6)
frame #32: _start + 0x25 (0x6576c5 in /usr/bin/python3)

Traceback (most recent call last):
  File "/app/train.py", line 54, in <module>
    run_training()
  File "/app/train.py", line 18, in run_training
    setup()
  File "/app/train.py", line 11, in setup
    dist.init_process_group(backend="nccl")
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/c10d_logger.py", line 83, in wrapper
    return func(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/c10d_logger.py", line 97, in wrapper
    func_return = func(*args, **kwargs)
                  ^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/distributed_c10d.py", line 1831, in init_process_group
    store, rank, world_size = next(rendezvous_iterator)
                              ^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/rendezvous.py", line 280, in _env_rendezvous_handler
    store = _create_c10d_store(
            ^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/rendezvous.py", line 190, in _create_c10d_store
    return TCPStore(
           ^^^^^^^^^
torch.distributed.DistNetworkError: The client socket has timed out after 600000ms while trying to connect to (2f3d44d8b0ea, 44403).
E0406 21:19:25.365000 131 torch/distributed/elastic/multiprocessing/api.py:986] failed (exitcode: 1) local_rank: 0 (pid: 161) of binary: /usr/bin/python3
I0406 21:19:25.376000 131 torch/distributed/elastic/multiprocessing/errors/__init__.py:375] ('local_rank %s FAILED with no error file. Decorate your entrypoint fn with @record for traceback info. See: https://pytorch.org/docs/stable/elastic/errors.html', 1)
Traceback (most recent call last):
  File "/usr/local/bin/torchrun", line 6, in <module>
    sys.exit(main())
             ^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 362, in wrapper
    return f(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/run.py", line 990, in main
    run(args)
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/run.py", line 981, in run
    elastic_launch(
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/launcher/api.py", line 170, in __call__
    return launch_agent(self._config, self._entrypoint, list(args))
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/launcher/api.py", line 317, in launch_agent
    raise ChildFailedError(
torch.distributed.elastic.multiprocessing.errors.ChildFailedError:
============================================================
train.py FAILED
------------------------------------------------------------
Failures:
  <NO_OTHER_FAILURES>
------------------------------------------------------------
Root Cause (first observed failure):
[0]:
  time      : 2026-04-06_21:19:25
  host      : dd3bc9f2214d
  rank      : 1 (local_rank: 0)
  exitcode  : 1 (pid: 161)
  error_file: <N/A>
  traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
============================================================
root@dd3bc9f2214d:/app# e
```

The reason is obvious.  The message `TCP client failed to connect/validate to host 2f3d44d8b0ea:44403` implies It tried to communicate via hostname, which is not reachable.

Looking at the first node’s log carefully, you see `master_addr=2f3d44d8b0ea`. This means the first node advertised itself with the hostname. In this case, we want to use the IP address.

Let’s find out where `torchrun` prints this log. Well, it’s in the log: `torch/distributed/elastic/agent/server/api.py:539`, [here](https://github.com/pytorch/pytorch/blob/v2.11.0/torch/distributed/elastic/agent/server/api.py#L539)​.  The line `master_addr = spec.master_addr or rdzv_info.bootstrap_store_info.master_addr` does that. It’s easy to confirm it with pdb.

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 torchrun   --nnodes=1 --nproc_per_node=1 --node_rank=0   --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500   --rdzv_conf=is_host=True   train.py
> /usr/local/bin/torchrun(6)<module>()
-> sys.argv[0] = sys.argv[0].removesuffix('.exe')
(Pdb) b /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/agent/server/api.py:524
Breakpoint 1 at /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/agent/server/api.py:524
(Pdb) c
I0406 21:51:17.139000 2208 torch/distributed/run.py:735] Using nproc_per_node=1.
I0406 21:51:17.236000 2208 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'is_host': 'True', 'timeout': 900}
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_qk97alc2
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 21:51:17.238000 2208 torch/distributed/launcher/api.py:224]
I0406 21:51:17.265000 2208 torch/distributed/elastic/agent/server/api.py:898] [default] starting workers for entrypoint: python3
I0406 21:51:17.267000 2208 torch/distributed/elastic/agent/server/api.py:693] [default] Rendezvous'ing worker group
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/agent/server/api.py(524)_rendezvous()
-> self._store = store
(Pdb) p spec.master_addr
None
(Pdb) p rdzv_info.bootstrap_store_info.master_addr
'2f3d44d8b0ea'
(Pdb)
```

It’s coming from `rdzv_info.bootstrap_store_info.master_addr`. Who set the hostname in it? We need to step in the line `spec.rdzv_handler.next_rendezvous()`.

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 torchrun   --nnodes=1 --nproc_per_node=1 --node_rank=0   --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500   --rdzv_conf=is_host=True   train.py
> /usr/local/bin/torchrun(6)<module>()
-> sys.argv[0] = sys.argv[0].removesuffix('.exe')
(Pdb) b /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/agent/server/api.py:514
Breakpoint 1 at /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/agent/server/api.py:514
(Pdb) c
I0406 23:42:36.311000 2320 torch/distributed/run.py:735] Using nproc_per_node=1.
I0406 23:42:36.408000 2320 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'is_host': 'True', 'timeout': 900}
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_fv9249b9
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 23:42:36.411000 2320 torch/distributed/launcher/api.py:224]
I0406 23:42:36.443000 2320 torch/distributed/elastic/agent/server/api.py:898] [default] starting workers for entrypoint: python3
I0406 23:42:36.446000 2320 torch/distributed/elastic/agent/server/api.py:693] [default] Rendezvous'ing worker group
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/agent/server/api.py(514)_rendezvous()
-> rdzv_info = spec.rdzv_handler.next_rendezvous()
(Pdb) s
--Call--
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py(1148)next_rendezvous()
-> def next_rendezvous(self) -> RendezvousInfo:
(Pdb) b /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1206
Breakpoint 2 at /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1206
(Pdb) c
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py(1206)next_rendezvous()
-> if self._bootstrap_store_info is None:
(Pdb) p self._bootstrap_store_info
None
(Pdb) self._this_node.addr
'2f3d44d8b0ea'
```

Alright, we have the hostname in `self._this_node.addr`, which is passed to `RendezvousStoreInfo.build` to build `_bootstrap_store_info`. Who set `self._this_node`? It’s in [`DynamicRendezvousHandler.__init__`](https://github.com/pytorch/pytorch/blob/v2.11.0/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py#L1086).

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 torchrun   --nnodes=1 --nproc_per_node=1 --node_rank=0   --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500   --rdzv_conf=is_host=True   train.py
> /usr/local/bin/torchrun(6)<module>()
-> sys.argv[0] = sys.argv[0].removesuffix('.exe')
(Pdb) b /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1072
Breakpoint 1 at /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1072
(Pdb) c
I0406 23:46:37.870000 2351 torch/distributed/run.py:735] Using nproc_per_node=1.
I0406 23:46:37.976000 2351 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'is_host': 'True', 'timeout': 900}
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_8d10hlwv
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 23:46:37.979000 2351 torch/distributed/launcher/api.py:224]
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py(1072)__init__()
-> if not settings.run_id:
(Pdb) p node
2f3d44d8b0ea_2351_0
(Pdb) p node.addr
'2f3d44d8b0ea'
(Pdb) bt
  /usr/local/bin/torchrun(7)<module>()
-> sys.exit(main())
  /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py(362)wrapper()
-> return f(*args, **kwargs)
  /usr/local/lib/python3.12/dist-packages/torch/distributed/run.py(990)main()
-> run(args)
  /usr/local/lib/python3.12/dist-packages/torch/distributed/run.py(981)run()
-> elastic_launch(
  /usr/local/lib/python3.12/dist-packages/torch/distributed/launcher/api.py(170)__call__()
-> return launch_agent(self._config, self._entrypoint, list(args))
  /usr/local/lib/python3.12/dist-packages/torch/distributed/launcher/api.py(284)launch_agent()
-> rdzv_handler=rdzv_registry.get_rendezvous_handler(rdzv_parameters),
  /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/registry.py(96)get_rendezvous_handler()
-> return handler_registry.create_handler(params)
  /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/api.py(377)create_handler()
-> handler = creator(params)
  /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/registry.py(48)_create_c10d_handler()
-> return create_handler(store, backend, params)
  /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py(1436)create_handler()
-> return DynamicRendezvousHandler.from_backend(
  /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py(1062)from_backend()
-> return cls(node, settings, backend.name, store, state_holder)
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py(1072)__init__()
-> if not settings.run_id:
(Pdb)
```

And `node` is coming from `cls._node_desc_generator.generate(local_addr)` in `from_backend`. Let’s see this `generate` function.

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 torchrun   --nnodes=1 --nproc_per_node=1 --node_rank=0   --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500   --rdzv_conf=is_host=True   train.py
> /usr/local/bin/torchrun(6)<module>()
-> sys.argv[0] = sys.argv[0].removesuffix('.exe')
(Pdb) b /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1049
Breakpoint 1 at /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1049
(Pdb) c
I0406 23:48:24.537000 2381 torch/distributed/run.py:735] Using nproc_per_node=1.
I0406 23:48:24.633000 2381 torch/distributed/launcher/api.py:131] Using default numa options = None
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'is_host': 'True', 'timeout': 900}
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_9e8y2vz6
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   numa_options             : None
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0406 23:48:24.635000 2381 torch/distributed/launcher/api.py:224]
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py(1049)from_backend()
-> node = cls._node_desc_generator.generate(local_addr)
(Pdb) p local_addr
None
(Pdb) s
--Call--
> /usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py(261)generate()
-> def generate(self, local_addr: str | None = None) -> _NodeDesc:
(Pdb)
```

`local_addr` is `None`, and see the function [`generate`](https://github.com/pytorch/pytorch/blob/v2.11.0/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py#L261). It does `local_addr or socket.getfqdn()`! This is where PyTorch prefers the hostname though the varialbe name is `local_addr`.

Now the question is how to overwrite it. In other words, how to set the IP address in this `local_addr`? We backtrack a little bit more. In `create_handler`, we pass `params.local_addr` to `DynamicRendezvousHandler.from_backend`. And `params` is `RendezvousParameters`. We already know this class, right? It’s where we store `is_host` earlier.

So we specify `--rdzv_conf=is_host=True,local_addr=10.100.10.1`?

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG STEPS=30 torchrun   --nnodes=1 --nproc_per_node=1 --node_rank=0   --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500   --rdzv_conf=is_host=True,local_addr=10.100.10.1   train.py
I0407 03:21:26.187000 2508 torch/distributed/run.py:735] Using nproc_per_node=1.
I0407 03:21:26.296000 2508 torch/distributed/launcher/api.py:131] Using default numa options = None
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   entrypoint               : train.py
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   min_nodes                : 1
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   max_nodes                : 1
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'is_host': 'True', 'local_addr': '10.100.10.1', 'timeout': 900}
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_zhdsgvr7
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   numa_options             : None
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0407 03:21:26.298000 2508 torch/distributed/launcher/api.py:224]
Traceback (most recent call last):
  File "/usr/local/bin/torchrun", line 6, in <module>
    sys.exit(main())
             ^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 362, in wrapper
    return f(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/run.py", line 990, in main
    run(args)
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/run.py", line 981, in run
    elastic_launch(
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/launcher/api.py", line 170, in __call__
    return launch_agent(self._config, self._entrypoint, list(args))
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/distributed/launcher/api.py", line 264, in launch_agent
    rdzv_parameters = RendezvousParameters(
                      ^^^^^^^^^^^^^^^^^^^^^
TypeError: torch.distributed.elastic.rendezvous.api.RendezvousParameters() got multiple values for keyword argument 'local_addr'
root@2f3d44d8b0ea:/app#
```

It was close, but failed with `TypeError: torch.distributed.elastic.rendezvous.api.RendezvousParameters() got multiple values for keyword argument 'local_addr'` because PyTorch instantiates `RendezvousParameters` as below. It explicitly specifies `local_addr` already, so we cannot include it in `config.rdzv_configs`.

```python
    rdzv_parameters = RendezvousParameters(
        backend=config.rdzv_backend,
        endpoint=config.rdzv_endpoint,
        run_id=config.run_id,
        min_nodes=config.min_nodes,
        max_nodes=config.max_nodes,
        local_addr=config.local_addr,
        **config.rdzv_configs,
    )
```

Where does `config.local_addr` come from? It’s from `args.local_addr` in `config_from_args`. `torchrun` directly takes `--local_addr` parameter.

Let’s try one more time! This is from the master node.

```log
root@2f3d44d8b0ea:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG NCCL_SOCKET_IFNAME=wg0 STEPS=30 torchrun   --nnodes=2 --nproc_per_node=1 --node_rank=0   --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500   --rdzv_conf=is_host=1   --local_addr=10.100.10.1   /app/train.py
I0407 03:33:29.820000 2911 torch/distributed/run.py:735] Using nproc_per_node=1.
I0407 03:33:29.926000 2911 torch/distributed/launcher/api.py:131] Using default numa options = None
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   entrypoint               : /app/train.py
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   min_nodes                : 2
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   max_nodes                : 2
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'is_host': '1', 'timeout': 900}
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_p9blggcp
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   numa_options             : None
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0407 03:33:29.928000 2911 torch/distributed/launcher/api.py:224]
I0407 03:33:29.956000 2911 torch/distributed/elastic/agent/server/api.py:898] [default] starting workers for entrypoint: python3
I0407 03:33:29.956000 2911 torch/distributed/elastic/agent/server/api.py:693] [default] Rendezvous'ing worker group
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/api.py:539] [default] Rendezvous complete for workers. Result:
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/api.py:539]   restart_count=0
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/api.py:539]   master_addr=10.100.10.1
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/api.py:539]   master_port=35121
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/api.py:539]   group_rank=0
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/api.py:539]   group_world_size=2
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/api.py:539]   local_ranks=[0]
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/api.py:539]   role_ranks=[0]
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/api.py:539]   global_ranks=[0]
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/api.py:539]   role_world_sizes=[2]
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/api.py:539]   global_world_sizes=[2]
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/api.py:539]   event_log_handler=null
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/api.py:539]
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/api.py:701] [default] Starting worker group
I0407 03:33:39.096000 2911 torch/distributed/elastic/agent/server/local_elastic_agent.py:299] use_agent_store: True
I0407 03:33:39.097000 2911 torch/distributed/elastic/agent/server/local_elastic_agent.py:195] Environment variable 'TORCHELASTIC_ENABLE_FILE_TIMER' not found. Do not start FileTimerServer.
I0407 03:33:39.097000 2911 torch/distributed/elastic/agent/server/local_elastic_agent.py:239] Environment variable 'TORCHELASTIC_HEALTH_CHECK_PORT' not found. Do not start health check.
2f3d44d8b0ea:2943:2943 [0] NCCL INFO ENV/Plugin: Could not find: libnccl-env.so
2f3d44d8b0ea:2943:2943 [0] NCCL INFO NCCL_SOCKET_IFNAME set by environment to wg0
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Bootstrap: Using wg0:10.100.10.1<0>
2f3d44d8b0ea:2943:2943 [0] NCCL INFO cudaDriverVersion 12080
2f3d44d8b0ea:2943:2943 [0] NCCL INFO NCCL version 2.28.9+cuda12.9
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Comm config Blocking set to 1
2f3d44d8b0ea:2943:2943 [0] NCCL INFO NET/Plugin: Could not find: libnccl-net.so
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Failed to open libibverbs.so[.1]
2f3d44d8b0ea:2943:2943 [0] NCCL INFO transport/net_ib.cc:852 -> 3
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Failed to initialize NET plugin IB
2f3d44d8b0ea:2943:2943 [0] NCCL INFO NCCL_SOCKET_IFNAME set by environment to wg0
2f3d44d8b0ea:2943:2943 [0] NCCL INFO NET/Socket : Using [0]wg0:10.100.10.1<0>
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Initialized NET plugin Socket
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Assigned NET plugin Socket to comm
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Using network Socket
2f3d44d8b0ea:2943:2943 [0] NCCL INFO ncclCommInitRankConfig comm 0x271b62c0 rank 0 nranks 2 cudaDev 0 nvmlDev 0 busId 70 commId 0x77d5d096c29cef94 - Init START
2f3d44d8b0ea:2943:2943 [0] NCCL INFO RAS client listening socket at ::1<28028>
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Bootstrap timings total 0.187971 (create 0.000071, send 0.000237, recv 0.183873, ring 0.000632, delay 0.000000)
2f3d44d8b0ea:2943:2943 [0] NCCL INFO NCCL_IGNORE_DISABLED_P2P set by environment to 1.
2f3d44d8b0ea:2943:2943 [0] NCCL INFO ncclTopoGetCpuAffinity: Affinity for GPU 0 is empty, ignoring. (GPU affinity =  ; CPU affinity = 0-27).
2f3d44d8b0ea:2943:2943 [0] NCCL INFO comm 0x271b62c0 rank 0 nRanks 2 nNodes 2 localRanks 1 localRank 0 MNNVL 0
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Channel 00/02 : 0 1
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Channel 01/02 : 0 1
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Trees [0] 1/-1/-1->0->-1 [1] -1/-1/-1->0->1
2f3d44d8b0ea:2943:2943 [0] NCCL INFO P2P Chunksize set to 131072
2f3d44d8b0ea:2943:2943 [0] NCCL INFO PROFILER/Plugin: Could not find: libnccl-profiler.so
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Check P2P Type isAllDirectP2p 1 directMode 0 isAllCudaP2p 1
2f3d44d8b0ea:2943:2977 [0] NCCL INFO [Proxy Service] Device 0 CPU core 5
2f3d44d8b0ea:2943:2978 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 24
2f3d44d8b0ea:2943:2943 [0] NCCL INFO TUNER/Plugin: Could not find: libnccl-tuner.so
2f3d44d8b0ea:2943:2943 [0] NCCL INFO threadThresholds 8/8/64 | 16/8/64 | 512 | 512
2f3d44d8b0ea:2943:2943 [0] NCCL INFO 2 coll channels, 2 collnet channels, 0 nvls channels, 2 p2p channels, 2 p2p channels per peer
2f3d44d8b0ea:2943:2943 [0] NCCL INFO CC Off, workFifoBytes 1048576
2f3d44d8b0ea:2943:2943 [0] NCCL INFO ncclCommInitRankConfig comm 0x271b62c0 rank 0 nranks 2 cudaDev 0 nvmlDev 0 busId 70 commId 0x77d5d096c29cef94 - Init COMPLETE
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Init timings - ncclCommInitRankConfig: rank 0 nranks 2 total 0.34 (kernels 0.13, alloc 0.00, bootstrap 0.19, allgathers 0.00, topo 0.01, graphs 0.00, connections 0.00, rest 0.00)
2f3d44d8b0ea:2943:2979 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 10
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Channel 00/0 : 1[0] -> 0[0] [receive] via NET/Socket/0
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Channel 01/0 : 1[0] -> 0[0] [receive] via NET/Socket/0
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Channel 00/0 : 0[0] -> 1[0] [send] via NET/Socket/0
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Channel 01/0 : 0[0] -> 1[0] [send] via NET/Socket/0
2f3d44d8b0ea:2943:2943 [0] NCCL INFO Connected all rings, use ring PXN 0 GDR 0
[Rank 0] Starting training...
Step 0 | Loss: 1.4985
Step 10 | Loss: 1.1718
Step 20 | Loss: 1.5557
[Rank 0] Training complete.
2f3d44d8b0ea:2943:2943 [0] NCCL INFO comm 0x271b62c0 rank 0 nranks 2 cudaDev 0 busId 70 - Destroy COMPLETE
2f3d44d8b0ea:2943:2943 [0] NCCL INFO ENV/Plugin: Closing env plugin ncclEnvDefault
I0407 03:33:43.109000 2911 torch/distributed/elastic/agent/server/api.py:917] [default] worker group successfully finished. Waiting 300 seconds for other agents to finish.
I0407 03:33:43.109000 2911 torch/distributed/elastic/agent/server/api.py:970] Local worker group finished (WorkerState.SUCCEEDED). Waiting 300 seconds for other agents to finish
I0407 03:33:43.119000 2911 torch/distributed/elastic/agent/server/api.py:984] Done waiting for other agents. Elapsed: 0.00843501091003418 seconds
```

And this is from the second node.  It all worked!

```log
root@dd3bc9f2214d:/app# NCCL_DEBUG=INFO LOGLEVEL=DEBUG NCCL_SOCKET_IFNAME=wg0 STEPS=30 torchrun   --nnodes=2 --nproc_per_node=1 --node_rank=1   --rdzv_id=test_job --rdzv_backend=c10d --rdzv_endpoint=10.100.10.1:29500   --local_addr=10.100.10.2   /app/train.py
I0407 03:33:37.836000 490 torch/distributed/run.py:735] Using nproc_per_node=1.
I0407 03:33:37.916000 490 torch/distributed/launcher/api.py:131] Using default numa options = None
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224] Starting elastic_operator with launch configs:
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   entrypoint               : /app/train.py
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   min_nodes                : 2
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   max_nodes                : 2
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   nproc_per_node           : 1
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   run_id                   : test_job
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   rdzv_backend             : c10d
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   rdzv_endpoint            : 10.100.10.1:29500
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   rdzv_configs             : {'timeout': 900}
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   max_restarts             : 0
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   monitor_interval         : 0.1
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   log_dir                  : /tmp/torchelastic_667h1_sh
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   metrics_cfg              : {}
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   event_log_handler        : null
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   numa_options             : None
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   signals_to_handle        : SIGTERM,SIGINT,SIGHUP,SIGQUIT
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   duplicate_stdout_filters : []
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]   duplicate_stderr_filters : []
I0407 03:33:37.917000 490 torch/distributed/launcher/api.py:224]
I0407 03:33:37.947000 490 torch/distributed/elastic/agent/server/api.py:898] [default] starting workers for entrypoint: python3
I0407 03:33:37.948000 490 torch/distributed/elastic/agent/server/api.py:693] [default] Rendezvous'ing worker group
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/api.py:539] [default] Rendezvous complete for workers. Result:
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/api.py:539]   restart_count=0
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/api.py:539]   master_addr=10.100.10.1
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/api.py:539]   master_port=35121
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/api.py:539]   group_rank=1
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/api.py:539]   group_world_size=2
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/api.py:539]   local_ranks=[0]
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/api.py:539]   role_ranks=[1]
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/api.py:539]   global_ranks=[1]
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/api.py:539]   role_world_sizes=[2]
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/api.py:539]   global_world_sizes=[2]
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/api.py:539]   event_log_handler=null
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/api.py:539]
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/api.py:701] [default] Starting worker group
I0407 03:33:39.097000 490 torch/distributed/elastic/agent/server/local_elastic_agent.py:299] use_agent_store: True
I0407 03:33:39.098000 490 torch/distributed/elastic/agent/server/local_elastic_agent.py:195] Environment variable 'TORCHELASTIC_ENABLE_FILE_TIMER' not found. Do not start FileTimerServer.
I0407 03:33:39.098000 490 torch/distributed/elastic/agent/server/local_elastic_agent.py:239] Environment variable 'TORCHELASTIC_HEALTH_CHECK_PORT' not found. Do not start health check.
dd3bc9f2214d:520:520 [0] NCCL INFO ENV/Plugin: Could not find: libnccl-env.so
dd3bc9f2214d:520:520 [0] NCCL INFO cudaDriverVersion 12080
dd3bc9f2214d:520:520 [0] NCCL INFO NCCL_SOCKET_IFNAME set by environment to wg0
dd3bc9f2214d:520:520 [0] NCCL INFO Bootstrap: Using wg0:10.100.10.2<0>
dd3bc9f2214d:520:520 [0] NCCL INFO NCCL version 2.28.9+cuda12.9
dd3bc9f2214d:520:520 [0] NCCL INFO Comm config Blocking set to 1
dd3bc9f2214d:520:520 [0] NCCL INFO NET/Plugin: Could not find: libnccl-net.so
dd3bc9f2214d:520:520 [0] NCCL INFO Failed to open libibverbs.so[.1]
dd3bc9f2214d:520:520 [0] NCCL INFO transport/net_ib.cc:852 -> 3
dd3bc9f2214d:520:520 [0] NCCL INFO Failed to initialize NET plugin IB
dd3bc9f2214d:520:520 [0] NCCL INFO NCCL_SOCKET_IFNAME set by environment to wg0
dd3bc9f2214d:520:520 [0] NCCL INFO NET/Socket : Using [0]wg0:10.100.10.2<0>
dd3bc9f2214d:520:520 [0] NCCL INFO Initialized NET plugin Socket
dd3bc9f2214d:520:520 [0] NCCL INFO Assigned NET plugin Socket to comm
dd3bc9f2214d:520:520 [0] NCCL INFO Using network Socket
dd3bc9f2214d:520:520 [0] NCCL INFO ncclCommInitRankConfig comm 0x26af6510 rank 1 nranks 2 cudaDev 0 nvmlDev 0 busId 70 commId 0x77d5d096c29cef94 - Init START
dd3bc9f2214d:520:520 [0] NCCL INFO RAS client listening socket at ::1<28028>
dd3bc9f2214d:520:520 [0] NCCL INFO Bootstrap timings total 0.007266 (create 0.000047, send 0.001406, recv 0.003567, ring 0.000768, delay 0.000001)
dd3bc9f2214d:520:520 [0] NCCL INFO NCCL_IGNORE_DISABLED_P2P set by environment to 1.
dd3bc9f2214d:520:520 [0] NCCL INFO ncclTopoGetCpuAffinity: Affinity for GPU 0 is empty, ignoring. (GPU affinity =  ; CPU affinity = 0-27).
dd3bc9f2214d:520:520 [0] NCCL INFO comm 0x26af6510 rank 1 nRanks 2 nNodes 2 localRanks 1 localRank 0 MNNVL 0
dd3bc9f2214d:520:520 [0] NCCL INFO Trees [0] -1/-1/-1->1->0 [1] 0/-1/-1->1->-1
dd3bc9f2214d:520:520 [0] NCCL INFO P2P Chunksize set to 131072
dd3bc9f2214d:520:520 [0] NCCL INFO PROFILER/Plugin: Could not find: libnccl-profiler.so
dd3bc9f2214d:520:520 [0] NCCL INFO Check P2P Type isAllDirectP2p 1 directMode 0 isAllCudaP2p 1
dd3bc9f2214d:520:553 [0] NCCL INFO [Proxy Service] Device 0 CPU core 17
dd3bc9f2214d:520:554 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 11
dd3bc9f2214d:520:520 [0] NCCL INFO TUNER/Plugin: Could not find: libnccl-tuner.so
dd3bc9f2214d:520:520 [0] NCCL INFO threadThresholds 8/8/64 | 16/8/64 | 512 | 512
dd3bc9f2214d:520:520 [0] NCCL INFO 2 coll channels, 2 collnet channels, 0 nvls channels, 2 p2p channels, 2 p2p channels per peer
dd3bc9f2214d:520:520 [0] NCCL INFO ncclCommInitRankConfig comm 0x26af6510 rank 1 nranks 2 cudaDev 0 nvmlDev 0 busId 70 commId 0x77d5d096c29cef94 - Init COMPLETE
dd3bc9f2214d:520:520 [0] NCCL INFO Init timings - ncclCommInitRankConfig: rank 1 nranks 2 total 0.17 (kernels 0.15, alloc 0.00, bootstrap 0.01, allgathers 0.00, topo 0.01, graphs 0.00, connections 0.00, rest 0.00)
dd3bc9f2214d:520:555 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 27
dd3bc9f2214d:520:520 [0] NCCL INFO Channel 00/0 : 0[0] -> 1[0] [receive] via NET/Socket/0
dd3bc9f2214d:520:520 [0] NCCL INFO Channel 01/0 : 0[0] -> 1[0] [receive] via NET/Socket/0
dd3bc9f2214d:520:520 [0] NCCL INFO Channel 00/0 : 1[0] -> 0[0] [send] via NET/Socket/0
dd3bc9f2214d:520:520 [0] NCCL INFO Channel 01/0 : 1[0] -> 0[0] [send] via NET/Socket/0
dd3bc9f2214d:520:520 [0] NCCL INFO Connected all rings, use ring PXN 0 GDR 0
[Rank 1] Starting training...
[Rank 1] Training complete.
dd3bc9f2214d:520:520 [0] NCCL INFO comm 0x26af6510 rank 1 nranks 2 cudaDev 0 busId 70 - Destroy COMPLETE
dd3bc9f2214d:520:520 [0] NCCL INFO ENV/Plugin: Closing env plugin ncclEnvDefault
I0407 03:33:43.118000 490 torch/distributed/elastic/agent/server/api.py:917] [default] worker group successfully finished. Waiting 300 seconds for other agents to finish.
I0407 03:33:43.118000 490 torch/distributed/elastic/agent/server/api.py:970] Local worker group finished (WorkerState.SUCCEEDED). Waiting 300 seconds for other agents to finish
I0407 03:33:43.120000 490 torch/distributed/elastic/agent/server/api.py:984] Done waiting for other agents. Elapsed: 0.0019180774688720703 seconds
```

### Conclusion

We just confirmed PyTorch DDP worked with a WireGuard network established between Docker containers.  There are special configurations needed:

* Specify `--rdzv_conf=is_host=1` for the master node because PyTorch doesn't see secondary IP addresses to check if the rendezvous endpoint is itself or not
* Specify `--local_addr` for every node to communicate via IP addresses instead of hostname

Actually specifying these parameters along with other standard DDP parameters such as `rdzv_endpoint` or `nnodes` is error-prone.  I created [a sample Docker image](https://github.com/kinesis-network/docker-image-samples/tree/main/11-torchrun) to semi-automated these configurations.

The question remains: Is this a bug in PyTorch we should report?

My answer is yes.  `--rdzv_conf=is_host=1` is a great workaround, but PyTorch should check all IP addresses assigned.  At the same time, I cannot find a clean solution for this yet.  It's strangely difficult to enumerate all IP addressese on Linux.  One solution AI suggested is to use [`fcntl.ioctl`](https://docs.python.org/3/library/fcntl.html#fcntl.ioctl) , but I believe it will be rejected because it looks too "C-style" or less compatible.  I'll think through it further.

Anyway, Kinesis now supports PyTorch DDP.  If you have a model to train, go to [https://portal.kinesis.network](https://portal.kinesis.network/) and run it!


# Why Compute Should Work Like Power on Tap

Imagine a world before the electricity grid. Every factory and household had to run its own generator — expensive, inefficient, and hard to scale. The breakthrough wasn’t just central power plants; it was the grid, the outlet, and the system that made electricity available everywhere, instantly. Compute is following the same trajectory. Today, organizations still buy, rent, and manage their own servers, or lock themselves into siloed cloud providers. At Kinesis, we believe the next leap is making compute work like electricity: seamless, reliable, and universally accessible. That’s the vision behind our electricity grid analogy.

## **The Electricity Grid Analogy**

#### **1. Generators → On-Prem Hardware**

In the early days, electricity was produced locally with generators. Every factory or household that wanted electricity had to generate its own.

💡 In compute, this is the same as enterprises buying servers and GPUs, running them in their own data centers, and putting them on the balance sheet.

***

#### **2. Power Plants → Cloud Providers**

Centralized power plants (coal, hydro, later nuclear) brought economies of scale, producing electricity more efficiently and selling it directly to customers.

💡 In compute, these are the hyperscalers (AWS, Azure, Google) and Tier 2 data centers. They run massive facilities, expose APIs, and sell compute capacity.

**Better economics than generators, but still siloed** — each provider has its own interfaces, contracts, and pricing models.

***

#### **3. The Grid (Wiring + Settlement) → Kinesis Protocol**

What really transformed electricity was the grid — the shared infrastructure that connected all power stations to all consumers.

The grid has two key elements:

* **Wiring (transmission lines):** the infrastructure that connects sources and destinations.
* **Settlement system/meters:** standardized measurement and billing, ensuring usage is recorded and suppliers are paid fairly.

💡In compute, this is the **Kinesis Protocol** (the Web3 layer):

* It connects all compute suppliers (hyperscalers, Tier 2 providers, crowdsourced GPUs, niche devices) to all customers.
* **Protocol + smart contracts = the meter and trust layer.** They measure usage, enforce performance guarantees, and handle settlement logic.
* **JOULE = the global compute currency.** Electricity grids are local and settle in local currencies; compute is inherently global. JOULE provides the first universal settlement unit — enabling per-second micropayments to millions of suppliers across borders.

***

#### **4. The Power Outlet → Kinesis Portal (IaaS)**

For consumers, the key innovation wasn’t the plant or the grid — it was the outlet. You just plug in, and electricity flows.

💡In compute, this is the **Kinesis Portal** (the IaaS interface):

* Customers upload workloads, set preferences, and simply “plug in.”
* All Web3 complexity (wallets, staking, tokens) is abstracted away.
* Enterprises pay in USD, suppliers can withdraw in USD — behind the scenes, every transaction is settled in JOULE.

***

#### **5. The Utility Company → Kinesis Network Inc. (KNI, Web2 Company)**

In electricity, your local utility manages your relationship with the grid. You pay them in your local currency, they ensure service continuity, and they handle wholesale settlement in the background.

💡In compute, this is **KNI (Kinesis Network Inc., the Web2 business):**

* Operates the customer-facing business.
* Sells compute via the Portal.
* Bills in USD, abstracts away protocol complexity, and guarantees SLAs.

***

#### **6. The Grid Operator → Kinesis Network Foundation (KNF, Web3 Company, governed by DAO)**

Electric grids are overseen by independent system operators (ISO) or regulators, responsible for long-term stability and expansion.

💡In compute, this is **KNF (Kinesis Network Foundation / DAO):**

* Owns and maintains the underlying protocol infrastructure.
* Issues and governs **JOULE**, the native currency of the protocol.
* Oversees the development, upgrade, and governance of the smart contracts that measure, enforce, and settle activity across the network.
* Its north star is to maximize the protocol’s long-term free cash flow and durability.
* Funds, supports, and accelerates ventures that increase protocol revenue.
* Over time, ensures **KNI is just one of many businesses plugged into the grid** — not the only one.

***

#### **7. The Appliances → Demand Side (Customers)**

Electricity became transformative when any device — from a light bulb to an entire factory — could plug into the same outlet and get exactly the power it needed. The outlet didn’t change, but what you could do with it scaled infinitely.

💡 In compute, this is the **demand side**:

* From **simple web apps and lightweight services** (like charging a phone or turning on a light bulb).
* To **AI inference workloads, ongoing enterprise applications, and smaller ML jobs** (like plugging in household appliances — steady, moderate power draw).
* To **large-scale AI training runs, scientific simulations, and industrial-grade HPC** (like running a steel mill, powering a subway system, or lighting up a skyscraper — massive continuous demand).

Customers all plug into the **same Kinesis Portal**. That single interface abstracts away the complexity of scale, delivering compute as flexibly as electricity itself — **a true utility, for every workload.**

***

The electricity grid turned power into a true utility — invisible, reliable, and available on demand. Compute is on the same path. By connecting fragmented supply, abstracting away complexity, and aligning incentives through a shared protocol, Kinesis makes compute flow as effortlessly as electricity from a wall outlet. Whether you’re running lightweight apps or training the largest AI models, the experience should be the same: plug in, and it just works. The future of compute isn’t about owning infrastructure — it’s about having compute on tap.


# Enable WireGuard on NVIDIA Jetson

By Toshihito Kikuchi

In [the earlier article](https://docs.kinesis.network/blog/another-day-of-debugging), I mentioned Dynamo supports ARM64 devices. At that time, I tested Amazon EC2 [T4g](https://aws.amazon.com/ec2/instance-types/t4/) Instances and [Raspberry Pi 5](https://www.raspberrypi.com/products/raspberry-pi-5/). As a third ARM64 device to test, I purchased [NVIDIA Jetson Orin Nano](https://developer.nvidia.com/embedded/learn/get-started-jetson-orin-nano-devkit).

After landing several patches, I was able to run Dynamo on Jetson. However, you need to do one extra step before installing Dynamo on your device: Build and install WireGuard on your end. Well, I could provide pre-built binaries, but I’d like to encourage people to build linux kernel by yourself. It’s fun. And Jetson as a Kinesis Node won’t be a mainstream scenario anyway.

I saw several forums discussing this topic too. Hopefully this article will help non-Kinesis users as well.

### Devices needed

I purchased:

1. NVIDIA Jetson Orin Nano Super Developer Kit ([Amazon](https://www.amazon.com/dp/B0BZJTQ5YP?ref=ppx_yo2ov_dt_b_fed_asin_title\&th=1))
2. Memory Card (I bought 64GB UHS-1 via [Amazon](https://www.amazon.com/dp/B0CYT352MW?ref=ppx_yo2ov_dt_b_fed_asin_title\&th=1). You pick any product.)

You also need a display supporting Display Port and a USB keyboard if you don’t have any.

### OS setup

Here's the official Getting Started Guide provided by NVIDIA:\
<https://developer.nvidia.com/embedded/learn/get-started-jetson-orin-nano-devkit>

In my case, the firmware was already up-to-date (36.4.x), so I directly downloaded SD Card Image of [JetPack 6.2.1](https://developer.nvidia.com/embedded/jetpack-sdk-621), but yours may not be. Don’t skip to check your firmware version. Please note that the latest version JetPack 7.0 is for Thor, not for Jetson Orin Nano (see <https://developer.nvidia.com/embedded/jetpack-archive>).

You can use your favorite writer to write an image to SD Card. I used [balena Etcher](https://etcher.balena.io/). Once it’s done, just insert your card and power on your Jetson. No drama here.

### Build WireGuard module

Let's double check WireGuard is not included in your device. You might be lucky to have it somehow. If it exists, it would be in `/lib/modules/$(uname -r)/kernel/drivers/net/wireguard`.

```
jetson@jetson:~$ lsb_release -a
No LSB modules are available.
Distributor ID: Ubuntu
Description:    Ubuntu 22.04.5 LTS
Release:        22.04
Codename:       jammy

jetson@jetson:~$ uname -a
Linux jetson 5.15.148-tegra #1 SMP PREEMPT Thu Sep 18 15:08:33 PDT 2025 aarch64 aarch64 aarch64 GNU/Linux

jetson@jetson:~$ ls -l /lib/modules/$(uname -r)/kernel/drivers/net
total 220
drwxr-xr-x  4 root root  4096 Oct  7 23:14 can
drwxr-xr-x  3 root root  4096 Jun 17 23:26 dsa
-rw-r--r--  1 root root 13265 Sep 18 15:40 dummy.ko
drwxr-xr-x 10 root root  4096 Jun 17 23:26 ethernet
drwxr-xr-x  2 root root  4096 Oct  7 23:14 ipa
-rw-r--r--  1 root root 39633 Sep 18 15:40 macvlan.ko
-rw-r--r--  1 root root 12177 Sep 18 15:40 macvtap.ko
drwxr-xr-x  2 root root  4096 Oct  7 23:14 mdio
-rw-r--r--  1 root root 10321 Sep 18 15:40 mdio.ko
drwxr-xr-x  2 root root  4096 Oct  7 23:14 pcs
drwxr-xr-x  2 root root  4096 Oct  7 23:14 phy
-rw-r--r--  1 root root 51305 Sep 18 15:40 tap.ko
drwxr-xr-x  2 root root  4096 Oct  7 23:14 usb
-rw-r--r--  1 root root 46089 Sep 18 15:40 veth.ko
drwxr-xr-x  2 root root  4096 Oct  7 23:14 vxlan
drwxr-xr-x  6 root root  4096 Oct  7 23:14 wireless
```

Since Linux 5.6, WireGuard is a [part](https://github.com/torvalds/linux/tree/master/drivers/net/wireguard) of kernel code in the mainstream, so you can get wireguard module by building linux kernel. Fortunately, NVIDIA provides [a guide](https://docs.nvidia.com/jetson/archives/r38.2/DeveloperGuide/SD/Kernel/KernelCustomization.html) on how to customize kernel.

The guide says we can download the kernel source with a script named “source\_sync.sh”. A weird thing is it doesn’t tell us where it is. It’s not included in the SD Card Image, JetPack. The answer is [here](https://forums.developer.nvidia.com/t/where-is-source-sync-sh/75149). It’s included in Driver Package. You can go to [Jetson Linux](https://developer.nvidia.com/embedded/jetson-linux) page and download "Driver Package (BSP)”. As of now, the package file is Jetson\_Linux\_r36.4.4\_aarch64.tbz2, which includes ./Linux\_for\_Tegra/source/source\_sync.sh.

To run source\_sync.sh, you need to specify “release-tag”, which the guide says is included in the release notes. But which release note should we check? AI tells me to do

```
jetson@jetson:/work$ head -n1 /etc/nv_tegra_release
# R36 (release), REVISION: 4.7, GCID: 42132812, BOARD: generic, EABI: aarch64, DATE: Thu Sep 18 22:54:44 UTC 2025
```

In the output above, the release is 36.4.7, but there is no corresponding release note. Given that the release note of [36.4.4](https://docs.nvidia.com/jetson/archives/r36.4.4/ReleaseNotes/Jetson_Linux_Release_Notes_r36.4.4.pdf) says its release tag is “jetson\_36.4.4”, we can make a guess. The source repo is <https://nv-tegra.nvidia.com/r/admin/repos/3rdparty/canonical/linux-jammy,tags>, so you can check the list of tags there.

```
jetson@jetson:/work/Linux_for_Tegra/source$ ./source_sync.sh -k -t jetson_36.4.7
Downloading default kernel/kernel-jammy-src source...
Cloning into '/work/Linux_for_Tegra/source/kernel/kernel-jammy-src'...
remote: Enumerating objects: 8617818, done.
remote: Counting objects: 100% (8617818/8617818), done.
remote: Compressing objects: 100% (1256098/1256098), done.
Receiving objects: 100% (8617818/8617818), 1.89 GiB | 14.65 MiB/s, done.
remote: Total 8617818 (delta 7322045), reused 8605598 (delta 7309852), pack-reused 0 (from 0)
Resolving deltas: 100% (7322045/7322045), done.
Checking objects: 100% (33554432/33554432), done.
The default kernel/kernel-jammy-src source is downloaded in: /work/Linux_for_Tegra/source/kernel/kernel-jammy-src
Syncing up with tag jetson_36.4.7...
Updating files: 100% (74252/74252), done.
Switched to a new branch 'mybranch_2025-10-08-1759908244'
/work/Linux_for_Tegra/source/kernel/kernel-jammy-src source sync'ed to tag jetson_36.4.7 successfully!
```

When we build linux kernel, we usually generate .config to customize build options with `menuconfig` or other tools. For Jetson Linux, however, the guide suggests to use `make -C kernel` to build, where Makefile applies [defconfig](https://nv-tegra.nvidia.com/r/plugins/gitiles/3rdparty/canonical/linux-jammy/+/refs/tags/jetson_36.4.7/arch/arm64/configs/defconfig) to build by default. Even if we run `menuconfig`, Makefile overrides .config with defconfig.

As you can see, `CONFIG_WIREGUARD` is not defined there.

```bash
jetson@jetson:/work/Linux_for_Tegra/source$ grep WIRE kernel/kernel-jammy-src/arch/arm64/configs/defconfig
CONFIG_SOUNDWIRE=m
CONFIG_SOUNDWIRE_QCOM=m
```

Let’s simply add `CONFIG_WIREGUARD=m` at the end of the file and kick off build. Once build is done, wireguard.ko is generated.

```bash
jetson@jetson:/work/Linux_for_Tegra/source$ tail -5 kernel/kernel-jammy-src/arch/arm64/configs/defconfig
# CONFIG_FUNCTION_GRAPH_TRACER is not set
# CONFIG_DYNAMIC_FTRACE is not set
CONFIG_MEMTEST=y

CONFIG_WIREGUARD=m

# Install prereq packages
jetson@jetson:/work/Linux_for_Tegra/source$ sudo apt install -y build-essential bc flex bison libssl-dev zstd libncurses
-dev

# Build! (Makefile specifies -j option accordingly)
jetson@jetson:/work/Linux_for_Tegra/source$ make -C kernel
...

jetson@jetson:/work/Linux_for_Tegra/source$ find . -name wireguard.ko
./kernel/kernel-jammy-src/drivers/net/wireguard/wireguard.ko
```

### Install WireGuard module

To install a single module, you simply copy a file under /lib/modules, generate dependencies, and load with `modprobe`.

```sh
sudo mkdir /lib/modules/$(uname -r)/kernel/drivers/net/wireguard/
sudo cp kernel/kernel-jammy-src/drivers/net/wireguard/wireguard.ko \
  /lib/modules/$(uname -r)/kernel/drivers/net/wireguard/
sudo depmod -a
sudo modprobe wireguard
```

Oops, the last `modprobe` command failed with the following error.

```bash
jetson@jetson:/work/Linux_for_Tegra/source$ sudo modprobe wireguard
modprobe: ERROR: could not insert 'wireguard': Unknown symbol in module, or unknown parameter (see dmesg)
```

Don’t panic. You can see more details with `dmesg`.

```bash
jetson@jetson:~$ sudo dmesg | tail
...
[10962.061383] wireguard: Unknown symbol chacha20poly1305_encrypt_sg_inplace (err -2)
[10962.061509] wireguard: Unknown symbol chacha20poly1305_encrypt (err -2)
[10962.061715] wireguard: Unknown symbol chacha20poly1305_decrypt_sg_inplace (err -2)
[10962.061798] wireguard: Unknown symbol xchacha20poly1305_encrypt (err -2)
[10962.061820] wireguard: Unknown symbol xchacha20poly1305_decrypt (err -2)
[10962.061913] wireguard: Unknown symbol chacha20poly1305_decrypt (err -2)
```

These are simple dependency errors. Searching code, you can easily find these symbols are implemented in libchacha20poly1305.ko, which is not installed in the kernel.

```
jetson@jetson:/work/Linux_for_Tegra/source/kernel/kernel-jammy-src/lib/crypto$ grep -nr chacha20poly1305_encrypt_sg_inplace
grep: libchacha20poly1305.o: binary file matches
chacha20poly1305-selftest.c:8928:		ret = chacha20poly1305_encrypt_sg_inplace(sg_src,
chacha20poly1305-selftest.c:9043:				if (!chacha20poly1305_encrypt_sg_inplace(sg_src,
grep: libchacha20poly1305.ko: binary file matches
grep: chacha20poly1305.o: binary file matches
chacha20poly1305.c:333:bool chacha20poly1305_encrypt_sg_inplace(struct scatterlist *src, size_t src_len,
chacha20poly1305.c:341:EXPORT_SYMBOL(chacha20poly1305_encrypt_sg_inplace);

jetson@jetson:/work/Linux_for_Tegra/source$ nm -g kernel/kernel-jammy-src/lib/crypto/libchacha20poly1305.ko | grep chacha20poly1305_encrypt_sg_inplace
0000000000000bd0 T chacha20poly1305_encrypt_sg_inplace
0000000037b34b92 A __crc_chacha20poly1305_encrypt_sg_inplace

jetson@jetson:/work/Linux_for_Tegra/source$ find /sys/module/$(uname -r) -name libchacha*
find: ‘/sys/module/5.15.148-tegra’: No such file or directory
```

Let’s install this module and try to load wireguard again.

```bash
sudo cp ./kernel/kernel-jammy-src/lib/crypto/libchacha20poly1305.ko \
  /lib/modules/$(uname -r)/kernel/lib/crypto/
sudo depmod -a
sudo modprobe wireguard
```

You’ll get the same error from `modprobe`, but it’s caused by different symbols.

```bash
[11355.067989] libchacha20poly1305: Unknown symbol poly1305_final_arch (err -2)
[11355.068039] libchacha20poly1305: Unknown symbol poly1305_init_arch (err -2)
[11355.068089] libchacha20poly1305: Unknown symbol poly1305_update_arch (err -2)
```

Repeat the same thing, searching code, identifying a missing module implementing the symbols. This time it’s poly1305-neon.ko. Neon is SIMD extension for ARM.

```bash
jetson@jetson:/work/Linux_for_Tegra/source/kernel/kernel-jammy-src/arch/arm64/crypto$ grep -nr poly1305_final_arch
grep: poly1305-neon.o: binary file matches
grep: poly1305-neon.ko: binary file matches
grep: poly1305-glue.o: binary file matches
poly1305-glue.c:170:void poly1305_final_arch(struct poly1305_desc_ctx *dctx, u8 *dst)
poly1305-glue.c:182:EXPORT_SYMBOL(poly1305_final_arch);
poly1305-glue.c:191:	poly1305_final_arch(dctx, dst);

jetson@jetson:/work/Linux_for_Tegra/source$ nm -g kernel/kernel-jammy-src/arch/arm64/crypto/poly1305-neon.ko | grep poly1305_final_arch
00000000f39f5240 A __crc_poly1305_final_arch
0000000000000e00 T poly1305_final_arch
```

Let’s just copy and try to load wireguard.ko once again.

```bash
sudo cp ./kernel/kernel-jammy-src/arch/arm64/crypto/poly1305-neon.ko \
  /lib/modules/$(uname -r)/kernel/arch/arm64/crypto/
sudo depmod -a
sudo modprobe wireguard
```

It should work this time, and you’ll see it’s loaded as below.

```bash
jetson@jetson:/work/Linux_for_Tegra/source$ lsmod | grep wire
wireguard              77824  0
libchacha20poly1305    16384  1 wireguard
ip6_udp_tunnel         20480  1 wireguard
udp_tunnel             24576  1 wireguard
libcurve25519_generic    36864  1 wireguard
ipv6                  471040  94 bridge,wireguard
```

### Verification

Dynamo uses wireguard through [our gateway container](https://hub.docker.com/r/kinesisorg/gateway-container). You can simulate this scenario by creating a minimum wireguard conf and running a container as follows. An endpoint can be random which doesn’t need to be valid for this verification purpose.

```bash
jetson@jetson:/work$ cat <<EOF > ${PWD}/wg-test.conf
[Interface]
PrivateKey = 6KvR4DZ1+uZom3S4VOBlbMtDpuZGGnbNEIFJfD4cjUg=
Address = 10.200.100.1/24
[Peer]
PublicKey = xFpe/+vmrUzQwtqpvVLB4AXygknLQo3f/qRR0AFXflQ=
AllowedIPs = 10.200.100.2/32
Endpoint = 1.2.3.4:51820
PersistentKeepalive = 25
EOF
jetson@jetson:/work$ docker run -d \
  --cap-add=NET_ADMIN \
  --device /dev/net/tun \
  --sysctl net.ipv4.ip_forward=1 \
  -v ${PWD}/wg-test.conf:/etc/wireguard/wg0.conf:ro \
  kinesisorg/gateway-container
250d35ac39c902515501ef863cf0d4cb91a2b407b75b6feab2fd6a41dc686c91
jetson@jetson:/work$ docker exec -it $(docker ps -q -f ancestor=kinesisorg/gateway-container) wg show
interface: wg0
  public key: i0JM5N5EVUDJ/08h5N4mt6bohSPrtXJZcNFqrsOULzA=
  private key: (hidden)
  listening port: 49594

peer: xFpe/+vmrUzQwtqpvVLB4AXygknLQo3f/qRR0AFXflQ=
  endpoint: 1.2.3.4:51820
  allowed ips: 10.200.100.2/32
  transfer: 0 B received, 592 B sent
  persistent keepalive: every 25 seconds
```

### Conclusion

Congratulations! Your Jetson device is ready to run Dynamo and join the Kinesis Network as a Global Node. In short, you need to clone Jetson Linux Kernel from NVIDIA’s repo, build it, and copy wireguard.ko along with a couple of dependent (crypto-related) modules. Pretty straightforward.

You might be interested in installing the kernel itself and debugging it. That’s good ambition. Let’s discuss it as another topic!


# Reaching out to home computers

By Toshihito Kikuchi

Kinesis Network transforms underutilized computing resources into scalable on-demand compute services. It’s expected most of them are home computers or on-premise corporate servers that don’t have public IP addresses. To achieve our goal, it’s essential to make these computers accessible from the Internet, but at the same time, we need to provide an unified, automated way to join the network. Asking owners of those devices to modify router settings or other network settings is not our option. In this article, I’m going to briefly explain how we deal with computers without having public IP addresses. By the way, we call these computers “Global Node”, in contract to “Datacenter Node”, to which a public IP address is assigned.

Whatever approach we take, we need to set up a globally accessible server that works as a proxy to connect to global nodes. In the earlier version of our product, we chose SSH and implemented a custom channel to forward each TCP connection. It worked well for a while, but we discontinued this approach because of scalability concern and high maintenance cost.

The next technology we chose is [WireGuard](https://www.wireguard.com/). We spotlighted this technology to solve another networking challenge, and then we realized this can be a replacement of SSH to achieve Global Nodes. WireGuard is basically a VPN. The basic idea is to install WireGuard on both the proxy side and the global node side and to redirect traffic to the proxy to the node through WireGuard network. In this setup, we call this proxy role “AppProxy”, in contrast to “NodeProxy”, which is to proxy a node’s admin channels (explained later).

A picture is worth a thousand words. Here’s the entire topology of a global node hosting an App with a single port.

<figure><img src="/files/jcGLj6NBdET6FOP3bzQ6" alt=""><figcaption></figcaption></figure>

In this diagram, you can see two outer boxes, Global Node and AppProxy. Both of them are Kinesis nodes, where Dynamo run and manage Kinesis Apps. Traffic comes from the right side. A node accepts two types of traffic: one is for normal app traffic like HTTP for Nginx, and the other is admin channels that host live docker log streams via gRPC and exec sessions via WebSocket. Admin channels are usually consumed by the Portal. An AppProxy opens accessible ports for these services, but on top of them, we have Load Balancer for App ports and NodeProxy for admin channels.

To establish a WireGuard network, we need to install WireGuard on both GlobalNode and AppProxy. Instead of directly installing WireGuard on these hosts, we decided to “kinesify” WireGuard and installed it as a Kinesis app \[1]. This design drastically simplifies our landscape because we can reuse most of code to set up WireGuard. We call a WireGuard on AppProxy “Frontend Gateway”, and the one on a Global Node “Backend Gateway”. As the diagram shows, we use 10.200.100.1 and 10.200.100.2 for all WireGuard networks. Using the same subnet for everything is to simplify our code. There is no specific reason why I chose this subnet. Maybe I should have chosen a more exotic subnet like 31.41.59.0/24. For docker bridge network, Dynamo selects an unused one among 192.168.0.0/24, 192.168.1.0/24, etc. This means a global node cannot run more than 255 apps (and one for admin channels). I think it’s a reasonable limit, but we can increase it by changing the network bits though it’s hardcoded.

Once a WireGuard network is established, a frontend gateway can access to a backend gateway via the WireGuard network’s IP even though the Global Node doesn’t have a global IP address.

The next step is to redirect traffic to AppProxy to the Global Node. We simply configure DNAT to redirect all traffic to AppProxy to the Global Node’s address on the WireGuard network. Here are the actual commands on AppProxy. To do this, we leverage a feature in Kinesis App to plant arbitrary files on a node and execute arbitrary commands on startup.

```
iptables -t nat -A PREROUTING -p tcp -j DNAT --to-destination 10.200.100.2
iptables -A FORWARD -p tcp -d 10.200.100.2 -j ACCEPT
iptables -t nat -A POSTROUTING -p tcp -d 10.200.100.2 -j MASQUERADE
```

For App ports, we configure one more DNAT to redirect traffic to the Global Node to the actual user app which is on the same Docker Bridge network. We create a custom network per App, so traffic to each App is completely separated on a global node.

Since admin channels are hosted by Dynamo Admin Server which is running on host, we cannot use DNAT unless we make host network accessible from a docker container, which is not recommended because of security concern. Therefore, we decided to include SSH in the gateway app and get Dynamo Admin Server to listen admin channels inside a backend gateway via [SSH remote port forwarding](https://en.wikipedia.org/wiki/Port_forwarding#Remote_port_forwarding).

That’s the overview of our global nodes. How does the network set up this whole thing? Here are the flow to establish admin channels when a global node joins the network:

1. Backend selects an AppProxy node to pair with
2. Backend installs a frontend gateway app on the AppProxy
3. Dynamo fetches dynamically assigned ports (WG’s listener, admin channels) and sends them back to the backend
4. Backend registers exposed admin channels with NodeProxy
5. Backend installs a backend gateway app on the global node with the config to connect to the AppProxy
6. Throughout the flow, Dynamo Admin Server keeps asking Dynamo for an SSH endpoint. Once it’s available, it automatically connects and listens ports inside a backend gateway.

When the network receives a request to run an App on a global node, it repeats almost the same steps to install frontend/backend gateway on both nodes.

\[1] Kinesified Gateway App: <https://github.com/kinesis-network/gateway-container>


# Another day of debugging

By Toshihito Kikuchi

One of the key components in the Kinesis ecosystem is Dynamo, an agent program to promote a computer to a computing node in the network. Recently, I made a patch to enable Dynamo to run on ARM64 devices. Usually we use Ubuntu as Linux distro for everyday use. On this day, however, I was in a mood to try something new and chose Amazon Linux in AWS. The story began there.

### Dynamo got stuck

On a node, we run Dynamo as a systemd service. When I started the Dynamo service on an Amazon Linux node, the process started, but I realized some important events were not logged. It seemed that the process was stuck in the middle of the startup phase.

The first triage is to run the exact same command manually, and when I ran the command on shell, it just worked normally. Here are the output: the former is from systemd; the latter is from manual run.

```jsx
#### systemd

[ec2-user@ip-172-31-15-164 kinesis-dynamo]$ sudo systemctl start dynamo
[ec2-user@ip-172-31-15-164 kinesis-dynamo]$ sudo journalctl -u dynamo -f
Aug 27 23:39:20 ip-172-31-15-164.ec2.internal systemd[1]: dynamo.service: Deactivated successfully.
Aug 27 23:39:20 ip-172-31-15-164.ec2.internal systemd[1]: Stopped dynamo.service - "Dynamo node service".
Aug 27 23:39:20 ip-172-31-15-164.ec2.internal systemd[1]: dynamo.service: Consumed 42.170s CPU time.
Aug 27 23:39:27 ip-172-31-15-164.ec2.internal systemd[1]: Started dynamo.service - "Dynamo node service".
Aug 27 23:39:27 ip-172-31-15-164.ec2.internal dynamo[37562]: time=2025-08-27T23:39:27.132Z level=INFO msg="Loaded your wallet" addr=0xa5e07b0a3944dd9158a8a72a6794004201669468 file=/opt/dynamo/id_ecdsa
Aug 27 23:39:27 ip-172-31-15-164.ec2.internal dynamo[37562]: time=2025-08-27T23:39:27.132Z level=INFO msg="Loaded AppCacheFile" file=/opt/dynamo/app-cache.json
Aug 27 23:39:27 ip-172-31-15-164.ec2.internal dynamo[37562]: time=2025-08-27T23:39:27.132Z level=INFO msg="Loaded a valid certificate" file=/opt/dynamo/backend.crt
Aug 27 23:39:27 ip-172-31-15-164.ec2.internal dynamo[37562]: time=2025-08-27T23:39:27.132Z level=INFO msg="Loaded config" file=/opt/dynamo/config.json
Aug 27 23:39:27 ip-172-31-15-164.ec2.internal dynamo[37562]: time=2025-08-27T23:39:27.182Z level=ERROR msg="Cannot initialize Docker Manager" err="Cannot connect to the Docker daemon at unix:///var/run/docker.sock. Is the docker daemon running?"
Aug 27 23:39:27 ip-172-31-15-164.ec2.internal dynamo[37562]: time=2025-08-27T23:39:27.183Z level=INFO msg="Serving gRPC" laddr=/tmp/kinesis-dynamo.sock
^C
[ec2-user@ip-172-31-15-164 kinesis-dynamo]$ ps -ef | grep node
root         778       2  0 22:38 ?        00:00:00 [xfs-inodegc/nvm]
ec2-user   37562       1  1 23:39 ?        00:00:00 /opt/dynamo/noded -config=/opt/dynamo/config.json
ec2-user   37577   36391  0 23:39 pts/0    00:00:00 grep --color=auto node

### manual run

[ec2-user@ip-172-31-15-164 kinesis-dynamo]$ /opt/dynamo/noded -config=/opt/dynamo/config.json
time=2025-08-27T23:40:03.620Z level=INFO msg="Loaded your wallet" addr=0xa5e07b0a3944dd9158a8a72a6794004201669468 file=/opt/dynamo/id_ecdsa
time=2025-08-27T23:40:03.620Z level=INFO msg="Loaded AppCacheFile" file=/opt/dynamo/app-cache.json
time=2025-08-27T23:40:03.621Z level=INFO msg="Loaded a valid certificate" file=/opt/dynamo/backend.crt
time=2025-08-27T23:40:03.621Z level=INFO msg="Loaded config" file=/opt/dynamo/config.json
time=2025-08-27T23:40:03.660Z level=ERROR msg="Cannot initialize Docker Manager" err="Cannot connect to the Docker daemon at unix:///var/run/docker.sock. Is the docker daemon running?"
time=2025-08-27T23:40:03.661Z level=INFO msg="Serving gRPC" laddr=/tmp/kinesis-dynamo.sock

!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
WARNING:

You should always run with libnvidia-ml.so that is installed with your
NVIDIA Display Driver. By default it's installed in /usr/lib and /usr/lib64.

!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
time=2025-08-27T23:40:03.767Z level=ERROR msg="Cannot initialize NVML" result=9
time=2025-08-27T23:40:03.802Z level=INFO msg="Successfully registered with Node Manager" address=0xa5e07b0a3944dd9158a8a72a6794004201669468 ip=""
^C
```

As you can see in the second output, we use [NVML](https://developer.nvidia.com/management-library-nvml) to collect information about NVIDIA GPU. The warning and the error of "Cannot initialize NVML” in the output are expected because this node doesn’t have any GPU. The problem is these warning and error were not logged from Dynamo running as a systemd service.

### Debugging with Delve

Since Dynamo is written in Go, the second triage is to attach [Delve](https://github.com/go-delve/delve) to a systemd process to debug. I could debug in interactive mode, but if the target process is a daemon like this case, I always start with batch mode i.e. running dlv to print all goroutines and dump it to a file. Why? If you pause the process for a long time, systemd may trigger restart. In general, you should keep your debug target fresh and untouched, like cooking fish. Dumping to a file also helps if the output is too long. You can use your favorite editor to analyze the output for unlimited time. Here are the commands to dump goroutines to a file.

```jsx
[ec2-user@ip-172-31-15-164 ~]$ sudo systemctl start dynamo
[ec2-user@ip-172-31-15-164 ~]$ ps -ef | grep noded
ec2-user   38720       1  3 23:54 ?        00:00:00 /opt/dynamo/noded -config=/opt/dynamo/config.json
ec2-user   38735   36391  0 23:54 pts/0    00:00:00 grep --color=auto noded
[ec2-user@ip-172-31-15-164 ~]$ dlv --allow-non-terminal-interactive attach 38720 <<< "goroutines -t 100" >goroutines.log
[ec2-user@ip-172-31-15-164 ~]$ ls -lh
total 64K
drwxr-xr-x.  4 ec2-user ec2-user  28 Aug 27 22:49 go
-rw-r--r--.  1 ec2-user ec2-user 45K Aug 27 23:54 goroutines.log
drwxr-xr-x. 22 ec2-user ec2-user 16K Aug 27 22:49 kinesis-dynamo
```

I found a groutine that looks indeed stuck.

```jsx
* Goroutine 36 - User: _cgo_gotypes.go:150 github.com/kinesis-network/kinesis-dynamo/nvml._Cfunc_NvmlInit (0x8fb5f0) (thread 38720) [select]
	 0  0x0000ffff87f68ff8 in [1m???[0m
	     at [1m?:-1[0m
	 1  0x000000000048de18 in [1mruntime.systemstack_switch[0m
	     at ./.go/src/runtime/[1masm_arm64.s:249[0m
	 2  0x000000000048572c in [1mruntime.cgocall[0m
	     at ./.go/src/runtime/[1mcgocall.go:185[0m
	 3  0x00000000008fb5f0 in [1mgithub.com/kinesis-network/kinesis-dynamo/nvml._Cfunc_NvmlInit[0m
	     at [1m_cgo_gotypes.go:150[0m
	 4  0x00000000008fb708 in [1mgithub.com/kinesis-network/kinesis-dynamo/nvml.Init.func1[0m
	     at ./kinesis-dynamo/nvml/[1mnvml.go:59[0m
	 5  0x00000000004998e0 in [1msync.(*Once).doSlow[0m
	     at ./.go/src/sync/[1monce.go:78[0m
	 6  0x00000000008fb6b4 in [1msync.(*Once).Do[0m
	     at ./.go/src/sync/[1monce.go:69[0m
	 7  0x00000000008fb6b4 in [1mgithub.com/kinesis-network/kinesis-dynamo/nvml.Init[0m
	     at ./kinesis-dynamo/nvml/[1mnvml.go:58[0m
	 8  0x00000000009bca08 in [1mgithub.com/kinesis-network/kinesis-dynamo/pulse.FillGpuInfo[0m
	     at ./kinesis-dynamo/pulse/[1mgpu.go:14[0m
	 9  0x0000000000d25f1c in [1mgithub.com/kinesis-network/kinesis-dynamo/core.(*Server).collectNodeInfo[0m
	     at ./kinesis-dynamo/core/[1mserver.go:389[0m
	10  0x0000000000d25720 in [1mgithub.com/kinesis-network/kinesis-dynamo/core.(*Server).advertiseSelf[0m
	     at ./kinesis-dynamo/core/[1mserver.go:340[0m
	11  0x0000000000d26c8c in [1mgithub.com/kinesis-network/kinesis-dynamo/core.(*Server).Serve.func1[0m
	     at ./kinesis-dynamo/core/[1mserver.go:464[0m
	12  0x0000000000490334 in [1mruntime.goexit[0m
	     at ./.go/src/runtime/[1masm_arm64.s:1268[0m
```

As mentioned above, Dynamo uses NVML. Since NVML is a C library (and Python bindings are available too), Dynamo implements a stub written in C to consume NVML APIs and compile/link it with [CGO](https://golang.google.cn/cmd/cgo/). The callstack above indicates that stub code never returns. Below is the actual Go code of Dynamo. The frame #7 in the output above indicates the call to `C.NvmlInit()` below never returns. Well, there is nothing interesting here in Go code.

```go
func Init(logger utils.ILogger) bool {
	nvmlInitOnce.Do(func() {
		if result := C.NvmlInit(); result == 0 {
			nvmlInitialized = true
			logger.Info("Initialized NVML")
		} else {
			logger.Error("Cannot initialize NVML", "result", result)
		}
	})
	return nvmlInitialized
}
```

What does `NvmlInit()` do? You may be surprised. Here’s the actual C code of Dynamo.

```c
nvmlReturn_t NvmlInit() {
  return nvmlInit();
}
```

It does nothing but calling a NVML function `nvmlInit` , which is actually an alias of a versioned function `nvmlInit_v2` . [NVIDIA’s reference](https://docs.nvidia.com/deploy/archive/R535/nvml-api/group__nvmlInitializationAndCleanup.html) has some description and notes about this method, but no information mentions a possible “stuck”.

So what’s next? NVML is closed-source unfortunately. Calling NVIDIA support?

### Debugging with GDB

Of course not. We dig deeper with [GDB](https://sourceware.org/gdb/) (Or you could go with [LLDB](https://lldb.llvm.org/) if you prefer). No source code is no problem.

Let’s restart dynamo and attach gdb to the process.

```c
[ec2-user@ip-172-31-15-164 ~]$ sudo systemctl restart dynamo
[ec2-user@ip-172-31-15-164 ~]$ ps -ef | grep noded
ec2-user   42723       1  2 01:52 ?        00:00:00 /opt/dynamo/noded -config=/opt/dynamo/config.json
ec2-user   42787   36391  0 01:52 pts/0    00:00:00 grep --color=auto noded
[ec2-user@ip-172-31-15-164 ~]$ gdb -q /opt/dynamo/noded
Reading symbols from /opt/dynamo/noded...
warning: File "/home/ec2-user/.go/.versions/1.25.0/src/runtime/runtime-gdb.py" auto-loading has been declined by your `auto-load safe-path' set to "$debugdir:$datadir/auto-load".
To enable execution of this file add
        add-auto-load-safe-path /home/ec2-user/.go/.versions/1.25.0/src/runtime/runtime-gdb.py
line to your configuration file "/home/ec2-user/.config/gdb/gdbinit".
To completely disable this security protection add
        set auto-load safe-path /
line to your configuration file "/home/ec2-user/.config/gdb/gdbinit".
For more information about this security protection see the
"Auto-loading safe path" section in the GDB manual.  E.g., run from the shell:
        info "(gdb)Auto-loading safe path"
(gdb) attach 42723
Attaching to program: /opt/dynamo/noded, process 42723
[New LWP 42788]
[New LWP 42773]
[New LWP 42759]
[New LWP 42739]
[New LWP 42729]
[New LWP 42728]
[New LWP 42727]
[New LWP 42726]
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/lib64/libthread_db.so.1".
runtime.futex () at /home/ec2-user/.go/src/runtime/sys_linux_arm64.s:651
651             SVC
Missing rpms, try: dnf --enablerepo='*debug*' install glibc-debuginfo-2.34-196.amzn2023.0.1.aarch64
```

If you want to dump callstacks of all threads as we did with dlv earlier, you can run

```c
(gdb) thread apply all bt
```

However, let’s try another approach this time. First, list all threads.

```c
(gdb) info threads
  Id   Target Id                                 Frame
* 1    Thread 0xffffbc4f5020 (LWP 42723) "noded" runtime.futex ()
    at /home/ec2-user/.go/src/runtime/sys_linux_arm64.s:651
  2    Thread 0xffff70f3e100 (LWP 42788) "noded" runtime.futex ()
    at /home/ec2-user/.go/src/runtime/sys_linux_arm64.s:651
  3    Thread 0xffff7194e100 (LWP 42773) "noded" runtime.futex ()
    at /home/ec2-user/.go/src/runtime/sys_linux_arm64.s:651
  4    Thread 0xffff7235e100 (LWP 42759) "noded" internal/runtime/syscall.Syscall6 ()
    at /home/ec2-user/.go/src/internal/runtime/syscall/asm_linux_arm64.s:17
  5    Thread 0xffff72e2e100 (LWP 42739) "noded" runtime.futex ()
    at /home/ec2-user/.go/src/runtime/sys_linux_arm64.s:651
  6    Thread 0xffff7395e100 (LWP 42729) "noded" runtime.futex ()
    at /home/ec2-user/.go/src/runtime/sys_linux_arm64.s:651
  7    Thread 0xffff743ae100 (LWP 42728) "noded" 0x0000ffffbc38aff8 in read () from /lib64/libc.so.6
  8    Thread 0xffff74dbe100 (LWP 42727) "noded" runtime.futex ()
    at /home/ec2-user/.go/src/runtime/sys_linux_arm64.s:651
  9    Thread 0xffff7580e100 (LWP 42726) "noded" runtime.futex ()
    at /home/ec2-user/.go/src/runtime/sys_linux_arm64.s:651
```

Probably we can ignore threads with `runtime.futex` because many are waiting there so it looks like Go’s management stuff. This means we’re interested in thread #4 or #7. Let’s check them one by one. In this case, thread #4 wasn’t interesting and #7 was what we were looking for.

```c
(gdb) thread 7
[Switching to thread 7 (Thread 0xffff62976100 (LWP 40398))]
#0  0x0000ffffaa952ff8 in read () from /lib64/libc.so.6
(gdb) bt
#0  0x0000ffffaa952ff8 in read () from /lib64/libc.so.6
#1  0x0000ffffaa8ed888 in __GI__IO_file_underflow () from /lib64/libc.so.6
#2  0x0000ffffaa8ee944 in _IO_default_uflow () from /lib64/libc.so.6
#3  0x0000ffffaa8e0ddc in _IO_getline_info () from /lib64/libc.so.6
#4  0x0000ffffaa8dfb20 in fgets () from /lib64/libc.so.6
#5  0x0000000000d824a4 in ?? ()
#6  0x0000000000d982c0 in nvmlInit_v2 ()
#7  0x0000000000d75c3c in NvmlInit () at nvml.c:87
#8  0x0000000000d757f0 in _cgo_4e2c63ea9b15_Cfunc_NvmlInit (v=0x40002259a8) at /tmp/go-build/cgo-gcc-prolog:114
#9  0x000000000049013c in runtime.asmcgocall () at /home/ec2-user/.go/src/runtime/asm_arm64.s:1049
#10 0x0000004000003dc0 in ?? ()
Backtrace stopped: not enough registers or memory available to unwind further
```

It’s in the middle of our stub function `NvmlInit` , which is calling `fgets` . Let’s confirm if `fgets` really never returns. This is an important step because this stuck could be caused by an infinite loop calling `fgets` repeatedly.

```c
(gdb) b *0xd824a4
Breakpoint 1 at 0xd824a4
(gdb) c
Continuing.
```

I left it for seconds and this breakpoint wasn’t hit. So Dynamo was stuck because `fgets` never returned. What kind of data are we reading from what? Standard input?

We’ll get there, but before that, let’s step back a little bit and remember this happens only when we run Dynamo via systemd. Two possibilities. If we run Dynamo directly from shell, `fgets` just returns without any problem, or `fgets` isn’t called at all. Let’s find it out. I took this approach because debugging a normal process is much easier than debugging a systemd process. We don’t have to check the PID, and we don’t have to worry about automatic restart, etc.

I just ran dynamo with gdb and set a breakpoint at `nvmlInit_v2`.

```c
[ec2-user@ip-172-31-15-164 ~]$ gdb -q /opt/dynamo/noded
Reading symbols from /opt/dynamo/noded...
warning: File "/home/ec2-user/.go/.versions/1.25.0/src/runtime/runtime-gdb.py" auto-loading has been declined by your `auto-load safe-path' set to "$debugdir:$datadir/auto-load".
To enable execution of this file add
        add-auto-load-safe-path /home/ec2-user/.go/.versions/1.25.0/src/runtime/runtime-gdb.py
line to your configuration file "/home/ec2-user/.config/gdb/gdbinit".
To completely disable this security protection add
        set auto-load safe-path /
line to your configuration file "/home/ec2-user/.config/gdb/gdbinit".
For more information about this security protection see the
"Auto-loading safe path" section in the GDB manual.  E.g., run from the shell:
        info "(gdb)Auto-loading safe path"
(gdb) b nvmlInit_v2
Breakpoint 1 at 0xd98264
(gdb) r -config=/opt/dynamo/config.json
Starting program: /opt/dynamo/noded -config=/opt/dynamo/config.json
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/lib64/libthread_db.so.1".
[New Thread 0xffffb130a100 (LWP 43272)]
[New Thread 0xffffabfff100 (LWP 43273)]
[New Thread 0xffffab5ef100 (LWP 43274)]
[New Thread 0xffffaabdf100 (LWP 43275)]
[New Thread 0xffffaa1cf100 (LWP 43276)]
time=2025-08-28T02:09:22.824Z level=INFO msg="Loaded your wallet" addr=0xa5e07b0a3944dd9158a8a72a6794004201669468 file=/opt/dynamo/id_ecdsa
time=2025-08-28T02:09:22.825Z level=INFO msg="Loaded AppCacheFile" file=/opt/dynamo/app-cache.json
time=2025-08-28T02:09:22.825Z level=INFO msg="Loaded a valid certificate" file=/opt/dynamo/backend.crt
time=2025-08-28T02:09:22.825Z level=INFO msg="Loaded config" file=/opt/dynamo/config.json
[New Thread 0xffffa97bf100 (LWP 43277)]
time=2025-08-28T02:09:22.873Z level=ERROR msg="Cannot initialize Docker Manager" err="Cannot connect to the Docker daemon at unix:///var/run/docker.sock. Is the docker daemon running?"
[New Thread 0xffffa8daf100 (LWP 43278)]
time=2025-08-28T02:09:22.874Z level=INFO msg="Serving gRPC" laddr=/tmp/kinesis-dynamo.sock
[Switching to Thread 0xffffabfff100 (LWP 43273)]

Thread 3 "noded" hit Breakpoint 1, 0x0000000000d98264 in nvmlInit_v2 ()
Missing rpms, try: dnf --enablerepo='*debug*' install glibc-debuginfo-2.34-196.amzn2023.0.1.aarch64
(gdb) bt
#0  0x0000000000d98264 in nvmlInit_v2 ()
#1  0x0000000000d75c3c in NvmlInit () at nvml.c:87
#2  0x0000000000d757f0 in _cgo_4e2c63ea9b15_Cfunc_NvmlInit (v=0x400038d9a8) at /tmp/go-build/cgo-gcc-prolog:114
#3  0x000000000049013c in runtime.asmcgocall () at /home/ec2-user/.go/src/runtime/asm_arm64.s:1049
#4  0x00000040000a4a80 in ?? ()
Backtrace stopped: not enough registers or memory available to unwind further
```

So far so good. Next, I set a breakpoint on `fgets` and continued.

```c
(gdb) b fgets
Breakpoint 2 at 0xfffff7e13a90
(gdb) c
Continuing.

!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
WARNING:

You should always run with libnvidia-ml.so that is installed with your
NVIDIA Display Driver. By default it's installed in /usr/lib and /usr/lib64.

!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
time=2025-08-28T02:09:33.691Z level=INFO msg="Disconnected from Runtime Manager" reason="connection error: desc = \\"error reading server preface: EOF\\""
[Detaching after vfork from child process 43280]

Thread 3 "noded" hit Breakpoint 2, 0x0000fffff7e13a90 in fgets () from /lib64/libc.so.6
(gdb) bt
#0  0x0000fffff7e13a90 in fgets () from /lib64/libc.so.6
#1  0x0000000000d82440 in ?? ()
#2  0x0000000000d982c0 in nvmlInit_v2 ()
#3  0x0000000000d75c3c in NvmlInit () at nvml.c:87
#4  0x0000000000d757f0 in _cgo_4e2c63ea9b15_Cfunc_NvmlInit (v=0x400038d9a8) at /tmp/go-build/cgo-gcc-prolog:114
#5  0x000000000049013c in runtime.asmcgocall () at /home/ec2-user/.go/src/runtime/asm_arm64.s:1049
#6  0x00000040000a4a80 in ?? ()
Backtrace stopped: not enough registers or memory available to unwind further
(gdb)
```

The breakpoint got hit nicely, but wait, did you notice the difference?

It’s obvious. We got NVML warning here. Remember we didn’t see this warning in the repro. So this call to `fgets` is not the one getting stuck. If you look at the stacktrace carefully, you’ll see the return address is different too. Here’s the callstack of the repro: d82440 vs d824a4.

```c
#4  0x0000ffffaa8dfb20 in fgets () from /lib64/libc.so.6
#5  0x0000000000d824a4 in ?? ()
#6  0x0000000000d982c0 in nvmlInit_v2 ()
#7  0x0000000000d75c3c in NvmlInit () at nvml.c:87
```

The imagebase might be different because this is a different process? That’s a good point, but probably it’s not the case. Usually the imagebase is reused for efficiency, but in this case the address is only 0x64 bytes different. If the module is mapped onto a different address, the difference would be aligned with page size 4k.

So this is a different call to `fgets`. Good news is both calls are made from the same function that is called from `nvmlInit_v2` . You can get it because the return address of a function calling `fgets` is 0xd982c0 in both cases.

Fun time to dive into assembly! Let’s start from `nvmlInit_v2`.

```c
(gdb) disas nvmlInit_v2
Dump of assembler code for function nvmlInit_v2:
   0x0000000000d98260 <+0>:     stp     x19, x30, [sp, #-16]!
   0x0000000000d98264 <+4>:     adrp    x19, 0x1aa1000
   0x0000000000d98268 <+8>:     bl      0xd825a0 <stubSpinLock>
   0x0000000000d9826c <+12>:    ldr     x0, [x19, #1080]
   0x0000000000d98270 <+16>:    cbz     x0, 0xd982a0 <nvmlInit_v2+64>
   0x0000000000d98274 <+20>:    adrp    x1, 0x1234000
   0x0000000000d98278 <+24>:    mov     w19, #0x0                       // #0
   0x0000000000d9827c <+28>:    add     x1, x1, #0x168
   0x0000000000d98280 <+32>:    bl      0x4103a0 <dlsym@plt>
   0x0000000000d98284 <+36>:    cbz     x0, 0xd98290 <nvmlInit_v2+48>
   0x0000000000d98288 <+40>:    blr     x0
   0x0000000000d9828c <+44>:    mov     w19, w0
   0x0000000000d98290 <+48>:    bl      0xd825d0 <stubUnLock>
   0x0000000000d98294 <+52>:    mov     w0, w19
   0x0000000000d98298 <+56>:    ldp     x19, x30, [sp], #16
   0x0000000000d9829c <+60>:    ret
   0x0000000000d982a0 <+64>:    bl      0xd824e0
   0x0000000000d982a4 <+68>:    str     x0, [x19, #1080]
   0x0000000000d982a8 <+72>:    cbnz    x0, 0xd98274 <nvmlInit_v2+20>
   0x0000000000d982ac <+76>:    adrp    x0, 0x1234000
   0x0000000000d982b0 <+80>:    add     x0, x0, #0x28
   0x0000000000d982b4 <+84>:    bl      0x410270 <puts@plt>
   0x0000000000d982b8 <+88>:    mov     w19, #0x9                       // #9
   0x0000000000d982bc <+92>:    bl      0xd823d0
   0x0000000000d982c0 <+96>:    bl      0xd825d0 <stubUnLock>
   0x0000000000d982c4 <+100>:   mov     w0, w19
   0x0000000000d982c8 <+104>:   ldp     x19, x30, [sp], #16
   0x0000000000d982cc <+108>:   ret
```

Right above 0xd982c0, there is a BL to 0xd823d0, which must be the function call we’re looking for.

One disadvantage of gdb compared to Windows Debugger (ntsd/cdb/kd/windbg) is the `disas` command (`uf` in Windows debuggers) doesn’t work if there is no matching symbol on the target address. We need to use `x` command instead.

```c
(gdb) x/80i 0xd823d0
   0xd823d0:    sub     sp, sp, #0x330
   0xd823d4:    mov     x2, #0x83                       // #131
   0xd823d8:    adrp    x1, 0x1233000 <crypto/internal/fips140/rsa..stmp_0+96>
   0xd823dc:    add     x1, x1, #0xe18
...snip...
   0xd82428:    bl      0x4103c0 <strstr@plt>
   0xd8242c:    cbnz    x0, 0xd824a8
   0xd82430:    mov     x2, x20
   0xd82434:    mov     x0, x19
   0xd82438:    mov     w1, #0x100                      // #256
   0xd8243c:    bl      0x410480 <fgets@plt>
   0xd82440:    cbnz    x0, 0xd82420
   0xd82444:    mov     x0, x20
   0xd82448:    bl      0x4103d0 <pclose@plt>
   0xd8244c:    bl      0x410130 <getpid@plt>
   0xd82450:    mov     w2, w0
   0xd82454:    adrp    x1, 0x1233000 <crypto/internal/fips140/rsa..stmp_0+96>
   0xd82458:    mov     x0, x22
   0xd8245c:    add     x1, x1, #0xeb0
   0xd82460:    bl      0x4100e0 <sprintf@plt>
   0xd82464:    add     x1, x23, #0xea0
   0xd82468:    mov     x0, x22
   0xd8246c:    bl      0x410160 <popen@plt>
   0xd82470:    mov     x20, x0
   0xd82474:    cbz     x0, 0xd824b0
   0xd82478:    adrp    x21, 0x1233000 <crypto/internal/fips140/rsa..stmp_0+96>
   0xd8247c:    add     x21, x21, #0xee0
   0xd82480:    b       0xd82494
   0xd82484:    mov     x1, x21
   0xd82488:    mov     x0, x19
   0xd8248c:    bl      0x4103c0 <strstr@plt>
   0xd82490:    cbnz    x0, 0xd824c4
   0xd82494:    mov     x2, x20
   0xd82498:    mov     x0, x19
   0xd8249c:    mov     w1, #0x100                      // #256
   0xd824a0:    bl      0x410480 <fgets@plt>            // <----- this causes stuck in the repro
   0xd824a4:    cbnz    x0, 0xd82484
...snip...
```

Alright, we found the call to `fgets` that is stuck in the repro. Well, we could enjoy finding out why it went down a different path with/without systemd, but not today. Let’s focus on the original issue.

The next thing to find out is what’s this `fgets` call. Before asking AI for help, just spend some time on reading this assembly code. The hint is what happens before we call `fgets`. Around there, we can see calls to standard functions like `getpid`, `sprintf`, `popen`, `strstr`, etc. If you have some experience in C programming, you can easily draw a picture about the behavior. Here’s my thinking process:

* There’s `popen` before `fgets` . Probably `fgets` is to read the output from a process spawned by this `popen`
* Before `popen`, the function calls `sprintf`. This must be to construct a command-line string to pass to `popen`
* `getpid` before `popen` is to embed the current process’s PID into the command-line string

What does this mean? Always step back and remember the original issue. The problem is `fgets` never returns. And we just found out `nvmlInit_v2` spawns a process and reads the output via `fgets` , which gets stuck. No documentation about this behavior in NVIDIA’s reference. Who knows a simple initialization function spawns a child process!

The next question is of course, what process is spawned? You can go back to the shell and run `ps` command to find out, but hold on, let’s stick to code and figure out the command line string first. `sprintf` is our next target because it looks like constructing a command to execute.

I believe you know `sprintf`. To understand a command to be constructed, it’s better to see a format string, which is passed to `sprintf` as the second argument. By the way, one advantage of ARM64 compared to x86 is fixed-sized instructions. Another advantage is a straightforward calling convention. The second argument is stored in the R1 register. Let’s see that value. Here’s the assembly to call `sprintf`.

```c
   0xd82450:    mov     w2, w0
   0xd82454:    adrp    x1, 0x1233000 <crypto/internal/fips140/rsa..stmp_0+96>
   0xd82458:    mov     x0, x22
   0xd8245c:    add     x1, x1, #0xeb0
   0xd82460:    bl      0x4100e0 <sprintf@plt>
```

If you’re not familiar with ARM64 instructions, specs can be found [here](https://developer.arm.com/documentation/dui0489/i/arm-and-thumb-instructions/). If you know ARM64 instructions already, you’ll easily see the second argument is `0x1233000 + 0xeb0`. Let’s check what’s there.

```c
(gdb) x/s 0x1233000 + 0xeb0
0x1233eb0:      "lsof -p %d | tr -s ' ' | cut -d ' ' -f 9 "
```

This is apparently a template of a command-line string, and explains why we call `getpid` .

Now, it’s time to confirm this finding. Start Dynamo and see if we’re really running this command.

```c
[ec2-user@ip-172-31-15-164 ~]$ sudo systemctl restart dynamo
[ec2-user@ip-172-31-15-164 ~]$ ps -ef | grep noded
ec2-user   45772       1  1 03:32 ?        00:00:00 /opt/dynamo/noded -config=/opt/dynamo/config.json
ec2-user   45789   36391  0 03:32 pts/0    00:00:00 grep --color=auto noded
[ec2-user@ip-172-31-15-164 ~]$ ps -ef | grep 45772
ec2-user   45772       1  0 03:32 ?        00:00:00 /opt/dynamo/noded -config=/opt/dynamo/config.json
ec2-user   45780   45772  0 03:32 ?        00:00:00 sh -c lsof -p 45772 | tr -s ' ' | cut -d ' ' -f 9
ec2-user   45781   45780  0 03:32 ?        00:00:00 lsof -p 45772
ec2-user   45784   45781 96 03:32 ?        00:00:09 lsof -p 45772
ec2-user   45791   36391  0 03:32 pts/0    00:00:00 grep --color=auto 45772

```

We’re running that command indeed! The next thing to do is of course, to run the exact same command manually.

```c
[ec2-user@ip-172-31-15-164 ~]$ sh -c lsof -p 45772 | tr -s ' ' | cut -d ' ' -f 9
SIZE/OFF
denied)
denied)
denied)
...snip...
/usr/lib64/libtirpc.so.3.0.0
/usr/lib64/libselinux.so.1
/usr/lib/ld-linux-aarch64.so.1
/usr/lib64/gconv/gconv-modules.cache
pipe
pipe
```

It just worked. Weird. Anyway, we know the root cause is not in Dynamo but in `lsof` , which doesn’t finish if it’s spawned from Dynamo spawned via systemd. What’s next? Should we attach gdb to a `lsof` process to find out? Well, I could, but at this time, I took a different approach, `strace`. And it worked very well.

### Root cause

Let’s restart dynamo and run `strace` for `lsof`.

```c
[ec2-user@ip-172-31-15-164 ~]$ sudo systemctl restart dynamo
[ec2-user@ip-172-31-15-164 ~]$ ps -ef | grep noded
ec2-user   46510       1  2 03:47 ?        00:00:00 /opt/dynamo/noded -config=/opt/dynamo/config.json
ec2-user   46526   36391  0 03:47 pts/0    00:00:00 grep --color=auto noded
[ec2-user@ip-172-31-15-164 ~]$ ps -ef | grep 46510
ec2-user   46510       1  0 03:47 ?        00:00:00 /opt/dynamo/noded -config=/opt/dynamo/config.json
ec2-user   46519   46510  0 03:47 ?        00:00:00 sh -c lsof -p 46510 | tr -s ' ' | cut -d ' ' -f 9
ec2-user   46520   46519  0 03:47 ?        00:00:00 lsof -p 46510
ec2-user   46524   46520 92 03:47 ?        00:00:09 lsof -p 46510
ec2-user   46528   36391  0 03:47 pts/0    00:00:00 grep --color=auto 46510
[ec2-user@ip-172-31-15-164 ~]$ strace -p 46520
strace: Process 46520 attached
wait4(-1, ^Cstrace: Process 46520 detached
 <detached ...>
[ec2-user@ip-172-31-15-164 ~]$ strace -p 46524
close(503905516)                        = -1 EBADF (Bad file descriptor)
close(503905517)                        = -1 EBADF (Bad file descriptor)
close(503905518)                        = -1 EBADF (Bad file descriptor)
close(503905519)                        = -1 EBADF (Bad file descriptor)
close(503905520)                        = -1 EBADF (Bad file descriptor)
close(503905521)                        = -1 EBADF (Bad file descriptor)
...
```

Hmm, we found something really unusual. In the `lsof` process, a bunch of `close` are failing.

So, it’s finally time to debug `lsof`? Again, I could, but to be more practical and save time, I thought it’s time to ask AI about what’s a possible cause. And here’s the actual answer I got.

```c
Yes — that strace output is a big hint.

When you see lsof looping on

```

close(503905516) = -1 EBADF close(503905517) = -1 EBADF …

```

it means lsof is walking a huge range of possible fd numbers and calling close() on them one by one.

That is not normal fd enumeration (which should only read /proc/<pid>/fd). This is an old implementation detail of lsof:

Before using /proc/<pid>/fd, lsof tries to "sanitize" the fd table by closing everything above its stdio fds, to make sure it doesn’t inherit descriptors it shouldn’t.

It doesn’t ask the kernel what the max fd is, but instead guesses from getrlimit(RLIMIT_NOFILE).
```

There is what I haven’t exposed yet: the service definition of Dynamo, which is this:

```c
[ec2-user@ip-172-31-15-164 ~]$ cat /etc/systemd/system/dynamo.service
[Unit]
Description="Dynamo node service"
After=network-online.target
StartLimitIntervalSec=0

[Service]
Type=simple
Restart=on-failure
RestartSec=5s
User=ec2-user
SyslogIdentifier=dynamo
WorkingDirectory=/opt/dynamo
KillSignal=SIGHUP
SyslogFacility=local7
ExecStart=/opt/dynamo/noded \\
  -config=/opt/dynamo/config.json
LimitNOFILE=infinity
LimitNPROC=infinity

[Install]
WantedBy=multi-user.target
```

Do you see anything unusual? It must be `LimitNOFILE=infinity`. There was no particular reason I added it. According to the AI’s answer, an old implementation of `lsof` guesses the max fd from `getrlimit(RLIMIT_NOFILE)` . In this case, it must be `RLIMIT_NOFILE = RLIM_INFINITY = 2^63-1` , meaning `lsof` is closing everything above its stdio fds below 2^63-1. This is the reason why Dynamo is stuck.

Alright, it seems that we found out the root cause, The next move is to confirm this. It’s easy. Let’s comment out `LimitNOFILE=infinity` and run Dynamo.

I don’t put the command output here because it’s obvious. The result is, it actually worked! Without `LimitNOFILE=infinity` , Dynamo just works normally.

What’s a conclusion here? It’s a bug in `lsof`? Basically yes, but remember what AI said: *This is an old implementation detail of lsof*. This implicates the latest `lsof` behaves differently. And remember I found this issue when I ran Dynamo on Amazon Linux ARM64, which recently started being supported. We’ve never seen this issue before, from x64 machines. Let’s check the version of `lsof`.

```c
[ec2-user@ip-172-31-15-164 ~]$ cat /proc/version
Linux version 6.1.147-172.266.amzn2023.aarch64 (mockbuild@ip-10-0-37-70) (gcc (GCC) 11.5.0 20240719 (Red Hat 11.5.0-5), GNU ld version 2.41-50.amzn2023.0.3) #1 SMP Thu Aug  7 19:28:45 UTC 2025
[ec2-user@ip-172-31-15-164 ~]$ lsof -v
lsof version information:
    revision: 4.94.0
    latest revision: <https://github.com/lsof-org/lsof>
    latest FAQ: <https://github.com/lsof-org/lsof/blob/master/00FAQ>
    latest (non-formatted) man page: <https://github.com/lsof-org/lsof/blob/master/Lsof.8>
    constructed: Mon May 19 00:00:00 UTC 2025
    compiler: cc
    compiler version: 11.5.0 20240719 (Red Hat 11.5.0-5) (GCC)
...
```

It’s 4.94.0. How about another machine of x64 where Dynamo runs without any problem?

```c
ubuntu@dynamo:~$ cat /proc/version
Linux version 6.8.0-59-generic (buildd@lcy02-amd64-035) (x86_64-linux-gnu-gcc-13 (Ubuntu 13.3.0-6ubuntu2~24.04) 13.3.0, GNU ld (GNU Binutils for Ubuntu) 2.42) #61-Ubuntu SMP PREEMPT_DYNAMIC Fri Apr 11 23:16:11 UTC 2025
ubuntu@dynamo:~$ lsof -v
lsof version information:
    revision: 4.95.0
    latest revision: <https://github.com/lsof-org/lsof>
    latest FAQ: <https://github.com/lsof-org/lsof/blob/master/00FAQ>
    latest (non-formatted) man page: <https://github.com/lsof-org/lsof/blob/master/Lsof.8>
    compiler: cc
    compiler version: 13.2.0 (Ubuntu 13.2.0-23ubuntu3)
```

Oops, it’s different: 4.94.0 vs 4.95.0!

Fortunately, unlike NVML, as you can see it in the output above, `lsof` is open-sourced. Actually [the release note of 4.95.0](https://github.com/lsof-org/lsof/releases/tag/4.95.0) mentions a fix for a bug that sound pretty much like our issue.

```
	[linux] use close_range instead of calling close repeatedly
	At the starting up, lsof closes its file descriptors greater
	than 2 by calling close(2) repeatedly. As reported in #186,
	it can take long time. Linux 5.9 introduced close_range(2).
	The new system call can close multiple file descriptors faster.
	@qianzhangyl reported the original issue (#186).
```

So I built `lsof` from source to see if the version 4.95.0 really fixed our issue (with uncommenting the line of `LimitNOFILE=infinity` of course). Unfortunately it didn’t. Thus I tried several version to find out which version fixed this issue. By the way, here are the commands to build `lsof`.

```c
$ git clone <https://github.com/lsof-org/lsof.git> -b 4.95.0
$ cd lsof
$ ./Configure linux
...
$ make -j4
...
$ ./lsof -v
lsof version information:
    revision: 4.95.0
    latest revision: <https://github.com/lsof-org/lsof>
    latest FAQ: <https://github.com/lsof-org/lsof/blob/master/00FAQ>
    latest (non-formatted) man page: <https://github.com/lsof-org/lsof/blob/master/Lsof.8>
    constructed by and on: ec2-user@ip-172-31-15-164.ec2.internal
    compiler: cc
    compiler version: 11.5.0 20240719 (Red Hat 11.5.0-5) (GCC)
...
```

Finally, I confirmed the issue is gone with 4.99.0 while the issue exists with 4.98.0. Let’s check [the release note of 4.99.0](https://github.com/lsof-org/lsof/releases/tag/4.99.0).

> ```
> 	[linux] Improve performance by using closefrom(). Closes #281.
> ```

The issue [281](https://github.com/lsof-org/lsof/issues/281) must be our issue. To really confirm, I could revert this patch and see, but I didn’t go that much further. Another mystery is why `lsof` 4.95.0 on Ubuntu x64 instance doesn’t hit this issue. That’s a very good question. I’ll dive into these another time.

Let’s recap:

* We saw an issue where Dynamo got stuck if it’s run through systemd on Amazon Linux ARM64.
* Dynamo calls `nvmlInit_v2`, which spawns `lsof` , which gets stuck
* lsof v4.94.0 has a bug that it attempts to close everything above its stdio fds
* Service definition of Dynamo specifies `LimitNOFILE=infinity` , meaning the upper limit of lsof’s close attempts is 2^63-1

Considering all the above, what’s a fix? I simply decided to specify `LimitNOFILE=1048576` instead of `infinity`. Again, there was no particular reason why I chose `infinity` from the beginning. Another possible approach is to apply this setting only if the version of `lsof` on the system is lower than 4.99.0, or set a custom value only while calling `nvmlInit_v2` , but I prefer the simplest fix.

This is just another day of debugging at Kinesis. Happy debugging!


# Biotech Is Ready to Save Lives. But Compute Is Holding It Back.

We stand on the threshold of one of the most profound leaps in human history.

AI is helping us decode the language of life itself. From folding proteins to designing molecules and mapping the causes of rare diseases, researchers are using the latest technological advances to accelerate scientific discovery and develop breakthrough treatments that were previously impossible to achieve. While the tools of modern biotech are changing medicine; they’re also changing what it means to be human.

But for the first time in history, progress toward these breakthroughs isn’t limited by scientific understanding. It’s limited by access to compute.

### The Future of Medicine Is Computational

A single human genome contains more than 3 billion base pairs. Decoding that data used to take years. Today, we can sequence a genome in hours, and AI is rapidly accelerating our ability to analyze and interpret it, often in days, sometimes minutes. But sequencing DNA is no longer the final goal. With AI, we're now asking these systems to interpret what the genetic code actually means.

Tools like AlphaFold and OpenFold have shown how neural networks can predict protein structures with remarkable accuracy, compressing decades of structural biology into hours of computation. AI-powered molecular simulations are speeding up early-stage drug discovery by screening candidate compounds at scale. And pattern recognition models are helping clinicians detect rare genetic disorders across vast, fragmented datasets, bringing new hope to patients who once went undiagnosed.

The era of precision medicine is just beginning, but its transformative potential hinges entirely on access to high-performance compute. This computational bottleneck affects researchers across the board. Even well-funded organizations [like the Chan Zuckerberg Initiative](https://www.businessinsider.com/mark-zuckerberg-priscilla-chan-gpus-recruitment-tool-chan-zuckerberg-initiative-2025-7) are prioritizing GPU access for their teams. This widespread demand for computing resources has created a fundamental challenge that threatens to limit the pace of medical breakthroughs.

### The Hidden Bottleneck No One Talks About

The cost of compute has become the silent constraint in modern science.

The same GPU clusters used to train language models are also needed to run biomedical AI. As demand skyrockets across every sector—from finance to entertainment—researchers are finding themselves priced out, queued up, or locked into rigid, expensive platforms.

Labs are postponing key analyses. Startups are redesigning their products around compute constraints. And sometimes, life-saving experiments are delayed indefinitely. This issue isn't because the science isn’t ready. The problem is the infrastructure isn’t ready.

When compute access becomes the bottleneck, human progress slows. When compute is wasted, live saving treatments are delayed.

### This Is a Global Problem. And It Needs a Global Solution.

Here’s the truth: the world doesn’t suffer from a lack of compute. It suffers from a failure to source and redistribute it to researchers at the forefront of medical research.

Across the planet, billions of dollars’ worth of CPUs and GPUs sit idle in data centers, offices, and even gaming consoles. Underutilized cloud inventory. Stranded enterprise servers. Consumer hardware with industrial-grade power. It’s all there—waiting.

What we lack is a unifying layer to make it usable.

Fixing problems at a global scale requires collaboration at a global scale. Scientists must be free to focus on discovery. Hardware providers should be able to contribute resources with minimal friction. What’s missing is the connective tissue: a new kind of digital infrastructure that makes compute accessible, affordable, and effortless to use—anywhere in the world.

### Kinesis Network: The Digital Infrastructure for Life-Saving AI

This is where Kinesis comes in.

Kinesis Network is building the critical digital backbone for AI-enabled biotech. We aggregate global compute capacity—from hyperscalers to edge devices—and route AI workloads in real time to the most available, cost-effective resources. Researchers simply upload their models and run.

No vendor lock-in. No orchestration headaches. No inflated bills.

Just pure, optimized performance—built for scientists, not sysadmins.

For the biotech ecosystem, this means:

* Running more simulations, faster.
* Scaling genomic analysis without hiring DevOps.
* Cutting compute costs by up to 99%, freeing budgets for research.

It means acceleration—without compromise.

### Everyone Can Play a Part

Here’s the remarkable part: you can be part of it too.

With Kinesis, anyone, anywhere in the world, can donate unused compute power to support the research they care about. Whether it’s cancer, Alzheimer’s, or rare genetic disorders, your idle machine can become part of a global mesh that fuels real scientific breakthroughs.

Just by running a lightweight application, your device can help power models that decode genomes, simulate drug interactions, or predict how a mutation might affect a protein. You don’t need to be a scientist to help cure disease. You just need a computer, and the willingness to contribute.

This is more than compute. It’s collaboration on lifesaving treatments at a planetary scale.

### A Shared Mission to Move Science Forward

Biotech has always been a long game. But today, time has never mattered more. Whether it's a child awaiting a rare disease diagnosis or a team racing to develop the next cancer therapy, every day of delay matters.

We believe that removing artificial bottlenecks from this process is one of the most meaningful contributions we can make to human progress. That’s why we’re building Kinesis. We are on a mission to support scientists, and accelerate the amazing work they are doing on behalf of the human species.

We don’t have the luxury to waste compute anymore.

Behind every delayed model is a patient.

Behind every slow simulation is a therapy that didn’t arrive in time.

And behind every bottleneck, there is an opportunity to do better.

Let’s fix this—together.


# DeepSeek & AI Compute: Disruption or Evolution?

DeepSeek’s latest AI breakthrough has stirred up conversations across the industry. Some call it a game-changer; others compare it to AI’s "Sputnik moment." The reaction has been swift—not just in AI circles but across the financial markets, with **over $600 billion wiped out from major tech stocks**, including **Nvidia**, **Microsoft**, **Alphabet**, and **Amazon**.

Nvidia, in particular, suffered **the biggest single-day loss in stock market history**, as investors reevaluated expectations about AI compute demand. The market panic stemmed from claims that DeepSeek’s model delivers comparable performance at a fraction of the compute cost, disrupting assumptions about the infrastructure requirements of cutting-edge AI.

But beyond the headlines and market volatility, what does this actually mean for AI infrastructure and compute efficiency?

Let’s break down what’s happening and what it means for the future of AI compute.

### **DeepSeek’s AI Training Efficiency: Impressive, But Not a Revolution**

DeepSeek trained its **671B parameter model using 2,048 GPUs over 57 days**, totaling **\~2.78 million GPU hours**. That’s an efficiency win compared to industry norms, but it doesn’t fundamentally change the compute landscape.

Key takeaways:

* **DeepSeek still required massive computational power.** Their process optimized GPU utilization, but did not eliminate the need for high-performance hardware.
* **This is an optimization, not a paradigm shift.** AI training remains an expensive, compute-heavy process, even with efficiency improvements.
* **DeepSeek leveraged Nvidia GPUs**—highlighting that current AI breakthroughs are still tied to the same core hardware ecosystem.

### **The Market Reacts: AI Disruption & the Compute Landscape Shift**

Following DeepSeek’s announcement, the stock market saw a sharp reaction, with over **$600 billion wiped out from major tech stocks** including Nvidia, Microsoft, Alphabet, and Amazon. Investors scrambled to reassess expectations about AI infrastructure needs and whether more efficient models could disrupt existing business models. However, history tells us that market shocks often overcorrect, and long-term compute demand remains resilient. AI adoption continues to expand, and while efficiency gains shift expectations, the need for scalable, cost-effective compute infrastructure is only growing.

* **More AI models, smarter architectures, and cost-efficient optimizations = Higher compute demand, not less.**
* **Nvidia will likely shift focus to inference acceleration.**
* **Decentralized compute solutions will become increasingly relevant.**

This is not a crisis for the compute industry—it’s an evolution.

### **The Real Cost in AI: Inference, Not Training**

While training is resource-intensive, inference is the **real long-term bottleneck**.

* Most AI companies do not train foundational models; they fine-tune existing ones.
* Training is typically a **one-time or infrequent cost**, whereas inference **scales linearly with usage**.
* Compute demand will continue to rise as AI adoption expands into real-world applications requiring continuous inference.

DeepSeek does not change this equation—it simply highlights that **optimization at every stage of AI compute is crucial.**

### **The Rise of Decentralized Compute**

One of the most overlooked impacts of AI efficiency improvements is the role of **decentralized compute networks**.

* **More efficient models fit better on smaller, distributed infrastructure.**
* **Gamer PCs, idle enterprise servers, and decentralized nodes** can now play a bigger role in AI processing.
* Hyperscalers will remain dominant, but **AI is no longer exclusive to massive centralized data centers.**

This shift is **great news for the entire AI ecosystem, including the open-source AI movement.** With DeepSeek making its model freely available (unlike GPT-4o), this raises new questions about the role of proprietary vs. open AI. Open-source AI could further fuel decentralized compute adoption, as companies look for cost-effective, flexible infrastructure alternatives to closed AI models —from startups building AI-driven products to enterprises integrating AI into their workflows, and decentralized compute providers enabling more efficient infrastructure. By reducing the reliance on hyperscalers and enabling more flexible compute access, this trend will lower costs and expand opportunities for AI adoption across industries.

### **Global AI Investment, Geopolitics & the Acceleration of Innovation**

DeepSeek proves one thing: **the AI race is heating up, and AI is now a geopolitical battleground.** The U.S. has imposed export controls on advanced AI chips to China, while China continues to make strides in AI research and alternative chip development. This competition will not only accelerate AI investments but also shape enterprise adoption strategies worldwide.

* China’s AI progress will likely accelerate **global AI investments**.
* The focus is shifting from raw compute power to **cost-efficient, scalable AI infrastructure.**

As AI compute becomes more efficient and widely accessible, innovation will accelerate. Startups will have lower barriers to entry, enterprises will be able to experiment with AI integrations more affordably, and decentralized compute networks will expand their role in supporting AI workloads. **This democratization of AI infrastructure will lead to new breakthroughs in model development**, fine-tuning, and application deployment.

### **What Comes Next? AI Compute Pricing & Sustainability**

The AI industry is moving toward a new phase where **efficiency, scalability, and decentralization** define success. But there’s another important factor—**AI compute pricing and sustainability.** If DeepSeek’s efficiency claims hold, will AI compute pricing come under pressure? Hyperscalers may respond by adjusting their pricing models, while decentralized compute providers could offer cost-competitive alternatives. Meanwhile, as AI models scale, energy consumption concerns grow—creating an opportunity for more sustainable, decentralized AI compute solutions. Key trends to watch:

* **LLMs are becoming commodities**—data ownership and enterprise adoption will be the real differentiators.&#x20;
* **Nvidia and other hardware players will double down on inference-focused chips.**
* **Decentralized compute networks will continue gaining traction** as AI models become more efficient.

### **Final Thoughts: The Future of AI Compute**

DeepSeek’s efficiency improvements are valuable, but they do not eliminate the need for massive compute power. Instead, they mark the beginning of a larger shift—one where geopolitics, open-source AI, pricing, and sustainability will shape the future of AI compute. They reinforce the importance of **optimization, inference efficiency, and scalable infrastructure**—areas where **decentralized compute can shine**.

At **Kinesis Network**, we’re building the future of AI compute—one that is **scalable, cost-effective, and decentralized.** As AI models evolve, so must the infrastructure that powers them. The future of AI isn’t just about who builds the biggest model—it’s about who can run them **the smartest.**

***

**Want to stay ahead of the AI compute revolution?** Follow [@Kinesis\_Network](https://twitter.com/kinesis_network) for more insights and updates on the future of compute.


# IEEE: The AI Boom Is Giving Rise to "GPU-as-a-Service

The industry harvests idle compute for AI startups that need it

The surge of interest in AI is creating a massive demand for computing power. Around the world, companies are trying to keep up with the vast amount of [GPUs](https://spectrum.ieee.org/tag/gpus) needed to power more and more advanced [AI models](https://spectrum.ieee.org/tag/ai-models). While [GPUs](https://spectrum.ieee.org/amd-mi300) are not the only option for [running an AI model](https://spectrum.ieee.org/ai-chip-sambanova), they have become the hardware of choice due to their ability to efficiently handle multiple operations simultaneously—a critical feature when developing [deep learning](https://spectrum.ieee.org/tag/deep-learning) models.

But not every AI startup has the capital to invest in the huge numbers of GPUs now required to run a cutting-edge model. For some, it’s a better deal to outsource it. This has led to the rise of a new business: GPU-as-a-Service (GPUaaS). In recent years, companies like [Hyperbolic](https://hyperbolic.xyz/), [Kinesis](https://kinesis.network/), [Runpod](https://www.runpod.io/), and [Vast.ai](https://vast.ai/) have sprouted up to remotely offer their clients the needed processing power.

While tech giants like [Amazon](https://spectrum.ieee.org/tag/amazon) or [Microsoft](https://spectrum.ieee.org/tag/microsoft) offering [cloud computing](https://spectrum.ieee.org/tag/cloud-computing) services own their infrastructure, smaller startups like Kinesis have created techniques to make the best out of the existing idle compute.

“*<mark style="color:green;">Businesses need compute. They need the model to be trained or their applications to be run; they don’t necessarily need to own or manage servers,</mark>*” says [Bina Khimani](https://www.linkedin.com/in/bkhimani/), co-founder of Kinesis.

[Studies](https://dl.acm.org/doi/10.1145/3597503.3639232) [have shown](https://www.nextplatform.com/2020/11/17/counting-the-cost-of-under-utilized-gpus-and-doing-something-about-it/) that more than half of the existing GPUs are not in use at any given time. Whether we’re talking personal computers or colossal server farms, a lot of processing capacity is under-utilized. What Kinesis does is identify idle compute—both for GPUs and CPUs—in servers worldwide and compile them into a single computing source for companies to use. Kinesis partners with universities, data centers, companies, and individuals who are willing to sell their unused computing power. Through a special software installed on their servers, Kinesis detects idle processing units, preps them, and offers them to their clients for temporary use.\
\
“*<mark style="color:green;">At Kinesis, we have developed technology to pool together fragmented, idle compute power and repurpose it into a server-less, auto-managed computing platform,</mark>*” says Khimani. Kinesis customers even have the possibility to choose from where they want their GPUs or CPUs to come.\
\
**AI Is Growing Faster Than Servers Can Keep Up**

GPUaaS is filling a growing gap in the AI industry. As learning models get more sophisticated, they need more power and an infrastructure that can process information faster and faster. In other words, without a sufficient number of GPUs, big AI models cannot operate—let alone improve. In October, OpenAI’s CEO, Sam Altman, [admitted](https://techcrunch.com/2024/10/31/openai-ceo-sam-altman-says-lack-of-compute-is-delaying-the-companys-products/) that the company was not releasing products as often as they had wished because they were facing “a lot of limitations” with their computing capacity.

Also in October, Microsoft’s CFO, Amy Woods, [told](https://www.microsoft.com/en-us/Investor/events/FY-2025/earnings-fy-2025-q1) the company’s investors in a conference call that demand for AI “*<mark style="color:green;">continues to be higher</mark>*” than their “*<mark style="color:green;">available capacity.</mark>*”

The biggest advantage of GPUaaS is economical. By removing the need to purchase and maintain the physical infrastructure, it allows companies to avoid investing in servers and IT management, and to instead put their resources toward improving their own deep learning, large language, and large vision models. It also lets customers pay for the exact amount of GPUs they use, saving the costs of the inevitable idle compute that would come with their own servers.

Server-less startups like Kinesis also claim to be friendlier to the environment than traditional cloud computing companies. By leveraging existing, unused processing units instead of powering additional servers, they say they significantly reduce energy consumption. In the last five years, big tech companies like [Google](https://spectrum.ieee.org/tag/google) and Microsoft have seen their [carbon emissions soar](https://www.npr.org/2024/07/12/g-s1-9545/ai-brings-soaring-emissions-for-google-and-microsoft-a-major-contributor-to-climate-change) due to the amount of energy consumed by AI. In response, some have turned their eyes to [nuclear energy](https://spectrum.ieee.org/nuclear-powered-data-center) to sustainably power their servers. Kinesis and other new startups offer a third route in which no further servers need to be plugged in.

“*<mark style="color:green;">Industry leaders are deeply committed to sustainability,</mark>*” Khimani says. “*<mark style="color:green;">With the focus on innovation and efficiency, they can optimize existing computing power that is already active and consuming energy, rather than continually adding more servers for every new application they run.</mark>*”

The growing demand for [machine learning](https://spectrum.ieee.org/tag/machine-learning) and colossal data consumption is turning GPUaaS into a very profitable tech sector. In 2023, the industry’s market size [was valued](https://www.fortunebusinessinsights.com/gpu-as-a-service-market-107797) at US $3.23 billion; in 2024, it grew to $4.31 billion. It’s expected to rise to $49.84 billion by 2032.\
\
“*<mark style="color:green;">The AI industry is rapidly advancing to a stage where the focus is shifting from merely building and training models to optimizing efficiency,</mark>*” Khimani says. “*<mark style="color:green;">Customers are increasingly asking questions like, ‘When training a new model, how can we do it extremely targeted and not consume an ocean of data that requires an enormous amount of compute and energy?’</mark>*”<br>


# Kinesis Network and Multiverse Computing Unite to Redefine AI Optimization

SEATTLE, WA, UNITED STATES, January 16, 2025

Amid headlines about corporations investing in nuclear power plants to meet the insatiable energy needs of AI workloads, two pioneering companies, [Multiverse Computing](https://multiversecomputing.com/) and [Kinesis Network Inc.](https://kinesis.network/) (“Kinesis”) have joined forces to provide a groundbreaking alternative—optimizing AI performance while drastically reducing resource consumption. The partnership comes as demand for AI capabilities grows exponentially, and organizations worldwide face mounting challenges related to power demands for these models.\
\
**Tackling Two Sides of the Same Problem**\
Multiverse Computing, a global leader in quantum and high-performance computing solutions, has been at the forefront of optimizing large language models (LLMs) to deliver improved performance and efficiency. By using advanced quantum-inspired algorithms, Multiverse Computing enhances the training and inference processes of LLMs, reducing computation time and energy requirements without compromising accuracy. The company’s compression software CompactifAI uses quantum-inspired tensor networks to improve the efficiency and performance of AI models. Shrinking these models reduces the compute power required to run them, which in turn reduces operating costs.\
\
“*<mark style="color:green;">We have found the perfect partner for our first significant U.S. partnership</mark>*<mark style="color:green;">,</mark>” said Enrique Lizaso Olmos, CEO and co-founder of Multiverse Computing. “*<mark style="color:green;">This collaboration will power new use cases for AI companies with models currently limited by power and compute requirements.</mark>*”\
\
Kinesis specializes in compute optimization, pooling distributed unused computing resources to maximize efficiency and minimize waste. Kinesis’ platform intelligently allocates workloads, ensuring that computational resources are used to their fullest potential. This dynamic resource optimization reduces infrastructure costs and aligns with global sustainability goals by decreasing carbon footprints.\
\
Kinesis’s innovative compute platform has also captured the attention of leading-edge organizations such as nCorium, Rare Compute, and Liminal Capital. By leveraging Kinesis’s optimization technology, these forward-thinking organizations aim to advance their sustainability journey while enhancing computational efficiency and reducing costs—without compromising performance. This growing interest from diverse industries highlights the versatility and transformative potential of Kinesis's approach to compute access and optimization.\
\
**A Collaborative Solution for Joint Customers**\
Together, Multiverse Computing and Kinesis are solving a pressing industry challenge: harnessing the power of AI while mitigating the environmental and financial costs associated with significant computational demands. By integrating their respective expertise in LLM and compute optimization, the two companies offer a comprehensive solution for customers seeking to drive innovation efficiently and sustainably.\
\
“<mark style="color:green;">AI workloads are pushing the limits of what current infrastructure can handle,</mark>” said Victor Gaspar, Chief Sales Officer of Multiverse Computing. “<mark style="color:green;">Our partnership with Kinesis allows us to offer a comprehensive approach to optimization, addressing both the performance of AI models and the utilization of compute resources. This collaboration ensures our customers can innovate without compromising on cost or sustainability.</mark>”\
\
Multiverse Computing recently completed the AWS Gen AI Accelerator program which included AWS credits, mentorship and resources to expand AI and machine learning (ML) technologies. Kinesis is an official AWS partner.\
\
“<mark style="color:green;">The synergy between Kinesis and Multiverse Computing is a game-changer for the AI industry,</mark>” said Bina Khimani, Chief Product and Revenue Officer of Kinesis. “<mark style="color:green;">By tackling inefficiencies in both model performance and computational resource allocation, we’re enabling enterprises to achieve more with less. It’s a powerful step forward for both business and the planet.</mark>”\
\
**Driving Efficiency and Sustainability**\
This partnership comes at a critical time when industries face increasing scrutiny over the environmental impact of AI technologies. By reducing energy consumption and optimizing resource usage, Multiverse Computing and Kinesis are driving operational efficiencies and setting a benchmark for sustainable AI practices.\
\
Together, Multiverse Computing and Kinesis are charting a course toward a more efficient and sustainable AI future, proving that innovation doesn’t have to come at the cost of the planet.\
\
**About Kinesis Network**\
Kinesis Network is a revolutionary platform that solves the critical shortage of computing access for AI developers, researchers, and enterprises. By connecting GPU providers with those who need computing power, Kinesis makes high-performance computing more efficient and accessible. Trusted by enterprises, academic institutions, and research organizations, Kinesis democratizes access to the vital infrastructure needed for AI development and data-intensive research. For more information about how Kinesis Network solves the compute access shortage, visit <https://kinesis.network/>.\
\
**About Multiverse Computing**\
Founded in 2019, Multiverse Computing is a leading quantum software company that combines quantum computing, AI and optimization solutions to solve complex industry challenges. The company’s team of over 160 full-time employees, comprising 40% PhDs and representing more than 43 nationalities, has developed CompactifAI, an LLM compressor which uses quantum-inspired tensor networks to make large language models (LLMs) more efficient and portable, reducing size by over 90%, with only a 2 – 3% drop in accuracy, and with over 50% savings in retraining and inference costs. Multiverse Computing enables organizations to optimize their AI models and computing resources for enhanced performance and efficiency. For more information, visit [www.multiversecomputing.com](http://www.multiversecomputing.com/).


# The Fundamentals of Web3: Revolutionizing the Internet with Ownership and Community

#### **What is Web3?**

Web3 is a multi-faceted term that encompasses a broad range of technologies and ideas aimed at decentralizing the internet. To better understand it, let’s look at how Web3 builds on its predecessors—Web1 and Web2—while advancing a new vision for how people interact online, with more user control over data, identity, and assets.

**Web1** was the first phase of the internet, characterized by static, read-only web pages. **Web2**, which followed, introduced interactive, dynamic experiences but brought about centralized control by large platforms. In contrast, **Web3** seeks to create a decentralized internet, where blockchain technology enables users to interact directly, without intermediaries, and with enhanced control over their data and digital presence.

Web3 envisions a trustless, permissionless internet where digital assets, data, and identities are directly controlled by individuals. This model is often described as the *ownership economy*, enabling users to own and control their data and digital assets through technologies like decentralized finance (DeFi) and non-fungible tokens (NFTs). Additionally, it promotes an interoperable and user-centric web, allowing for seamless interactions across platforms while empowering users to participate in decision-making processes via decentralized governance.

***

#### **The Evolution of the Web**

1. **Web1** – The “Read-Only” Web: Early websites displayed static content that users could only read or view.
2. **Web2** – The “Read-Write” Web: Enabled dynamic content and interactivity, which allowed users to participate in social networks, online marketplaces, and more.
3. **Web3** – The “Read-Write-Own” Web: Empowers users to own, control, and monetize their data, assets, and digital presence.

By building on blockchain and decentralized technologies, Web3 has the potential to redefine industries and reshape the internet as we know it.

***

#### **Core Principles of Web3**

1. **Decentralization**: Unlike Web2 platforms that are owned and operated by central entities, Web3 applications are often decentralized, distributing control across participants rather than a single authority.
2. **User Sovereignty**: Web3 gives users control over their digital assets, data, and identity. This is often achieved through cryptographic wallets, where only the user has access.
3. **Trustless Transactions**: Transactions on Web3 platforms don’t require third-party intermediaries. Instead, smart contracts—self-executing agreements on the blockchain—enable secure peer-to-peer interactions.
4. **Interoperability**: Web3 promotes interoperability, meaning different applications and platforms can connect and share information seamlessly. This allows users to move digital assets and data across ecosystems without friction.
5. **Transparency**: Since blockchain transactions are publicly recorded, Web3 applications are inherently transparent, allowing users to verify the validity of transactions and governance activities.

***

#### **Key Sectors in Web3**

Web3 spans a range of sectors, each with distinct applications and use cases. Here are a few examples that highlight the diversity of Web3’s potential:

* **DeFi (Decentralized Finance)**: Offers financial services without intermediaries, allowing for peer-to-peer lending, borrowing, and trading on decentralized platforms.
* **DePIN (Decentralized Physical Infrastructure Networks)**: Powers real-world infrastructure like telecommunications, cloud storage, and compute networks by incentivizing resource sharing.
* **NFTs (Non-Fungible Tokens)**: Enables verifiable ownership of digital assets such as art, music, collectibles, and real estate in the digital space.
* **Social & Creator Economies**: Empowers creators to monetize their work directly through tokenized communities, where fans and followers gain ownership in the creator's success.

These sectors highlight Web3’s versatility, reshaping traditional industries from finance and media to infrastructure and beyond.

***

#### **Community-Driven Funding and Ownership**

One of the defining features of Web3 is its approach to funding and ownership. Unlike Web2 companies, which often rely on centralized funding from venture capital, Web3 projects tend to involve the community early on. This community-first model allows for early user participation, giving community members a vested interest in a project’s success and fostering a collaborative environment.

**Utility tokens** are frequently used in Web3 to provide holders access to specific functions within the ecosystem, rather than being a form of ownership. These tokens often serve as governance rights, staking assets, or utility credits, allowing holders to interact with the network while aligning with the project’s purpose. For instance, community staking or participation-based incentives encourage early adopters to actively engage with the project without relying on speculative token sales. This decentralized approach to funding and ownership enables community members to directly support the project’s growth, fostering a more equitable ecosystem.

***

#### **New Business Models in Web3**

In addition to technical advancements, Web3 introduces innovative business models that change how value is created and shared:

* **Token Economies**: Tokenized ecosystems incentivize participation by rewarding users with tokens for their contributions.
* **DAOs (Decentralized Autonomous Organizations)**: Enable community governance by allowing token holders to vote on key decisions, making the community part-owners of the project.
* **Play-to-Earn and Participate-to-Earn**: These models reward users for engaging with the ecosystem, be it through gaming, community participation, or content creation.
* **Data Sovereignty**: Web3 allows users to retain control over their personal data, with some models allowing users to earn from their own data.

These models challenge traditional approaches by prioritizing community involvement, shared ownership, and equitable value distribution, offering a fresh approach to how digital platforms operate.

***

#### **Looking Ahead**

Web3 represents a paradigm shift that goes beyond mere technological advancements. It fosters a new relationship between users, platforms, and digital assets, empowering individuals to take control of their data and assets while participating in decentralized economies. As Web3 continues to evolve, it will likely bring about a wave of innovation across sectors and industries, presenting new opportunities and challenges. Embracing Web3’s principles offers a glimpse into a future internet that is more decentralized, transparent, and community-oriented, with significant implications for society and the global economy.


# Kinesis Network Saves Costs

In recent years, serverless computing has emerged as a transformative approach to running applications and workloads in the cloud. This model shifts responsibility for infrastructure provisioning and management from the end user to the cloud provider. In this post, we are exploring how Kinesis Network with its serverless architecture saves costs.&#x20;

**1. Over-Provisioning and Idle Capacity**

One of the most notable causes of waste in an instance-based computing model is over-provisioning. When organizations provision specific virtual machines or containers, they frequently allocate more resources—CPU, memory, or storage—than the application actively needs. This over-allocation often aims to ensure enough capacity is available for peak usage, even if those peaks only occur sporadically. Unfortunately, during low-traffic intervals, these resources remain largely underutilized, creating wasted capacity. Our experience has shown us that a large portion of instances see less than 20% utilization on average, sometimes as bad as only 1%. In contrast, a serverless environment scales automatically based on demand. **The Kinesis Network spins up or tears down resources as needed, eliminating over-provisioning and shrinking the idle footprint significantly.**

**2. Pay-For-Idle vs. Pay-Per-Use**

Hand-in-hand with over-provisioning is the pay-for-idle model inherent in instance-based services. Organizations pay for entire virtual machines regardless of whether they are actively processing tasks or sitting idle. This can be a significant cost drain for applications with unpredictable traffic or infrequent usage patterns. **By contrast, Kinesis services follow a pay-per-use billing model.** This means that costs accrue only when the application processes requests. With no fixed cost for idle time, developers can benefit from substantial cost savings, making serverless the more financially efficient option.

**3. Operational Complexity**

In an instance-based model, DevOps teams are responsible for provisioning, configuring, patching, and maintaining virtual machines. This operational complexity often leads to less predictable results and potential inefficiencies. Each instance must be monitored, and scaling must be managed carefully. This level of manual oversight not only consumes time and resources but also increases the likelihood of human error. In our serverless architecture, **Kinesis Network handles most of these operational responsibilities—provisioning, scaling, fault tolerance, and more.** The result is less overhead in terms of both personnel and budget, leading to a more focused environment where developers can prioritize core business logic.

**4. Environmental Impact**

The wasteful nature of running underutilized or idle virtual machines has broader implications beyond cost. Modern data centers require substantial energy resources to power servers and maintain cooling systems. When capacity is over-allocated, these underutilized servers continue to consume power, leaving a significant carbon footprint. Kinesis Network, with its on-demand approach, reduces total operating hours of hardware. **Because Kinesis Network allocate resources dynamically, computing resources remain dormant until needed, cutting down on energy usage and subsequent environmental impact.**

**5. Flexibility and Agility**

From a software development perspective, serverless computing enables agile development. Functions can be deployed quickly without detailed infrastructure management, allowing teams to iterate faster. In an instance-based setup, teams must often navigate lengthy processes to provision and configure additional instances or adjust the size of existing ones. This overhead can stifle innovation and extend deployment cycles. **In contrast, Kinesis Network streamlines these processes, ensuring teams can rapidly experiment, test, and roll out new features with fewer constraints—and less waste.**

***

Despite these advantages, it's important to note that serverless computing isn't a silver bullet for all workloads. Long-running processes or applications with consistent, predictable loads might still benefit from instance-based deployments. However, for the vast majority of modern applications with variable workloads, serverless architectures offer a more environmentally sustainable approach to cloud computing.


# Powerlifting with Kinesis Network

Kinesis Network is designed to be flexible and multi-purpose. However, some workloads stand out to benefit most from the capabilities of Kinesis Network: Applications that  often rely on **parallel computation** (splitting large tasks into smaller ones) and **optimized hardware acceleration.**

Below are some common examples of workloads or applications that tend to push CPUs and GPUs to their limits:

### 1. AI & Machine Learning

#### Neural Network Training

* **What It Is**: Training large-scale models (e.g., convolutional neural networks for image recognition, transformers for language tasks) involves iterating over massive datasets and performing billions of floating-point operations.
* **Why It’s Intensive**: Each forward and backward pass can update millions or even billions of parameters. GPUs excel at this kind of parallelizable matrix multiplication.
* **Examples**:
  * **Image Classification** (ResNet, EfficientNet) on platforms like ImageNet.
  * **Large Language Models** (GPT-style), which can have hundreds of billions of parameters.
  * **Recommendation Systems** at companies like Netflix or YouTube, which process user activity logs in real time to update models.

#### Inference & Real-Time Prediction

* **What It Is**: Once models are trained, they need to make predictions on new data quickly.
* **Why It’s Intensive**: High-traffic systems (like voice assistants, search engines, or real-time translation services) can receive millions of queries per second. Optimizing inference—often on GPUs or specialized hardware (FPGAs, TPUs)—is critical for low latency.
* **Examples**:
  * **Virtual Assistants** (Possible alternatives to Siri, Alexa, Google Assistant).
  * **Security Systems** (Real time analysis of security footage, pattern matching)

#### Reinforcement Learning & Robotics

* **What It Is**: Training agents to interact with environments (e.g., game playing, industrial robots).
* **Why It’s Intensive**: Simulation-based training (like AlphaGo/AlphaZero) can require playing millions of matches or environment steps.
* **Examples**:
  * **Gaming engines like Go or Chess** playing against themselves.
  * **Industrial Robotics** where digital twins simulate thousands of robotic arm movements to optimize tasks.

***

### 2. Scientific Simulations & High-Performance Computing (HPC)

#### Weather Forecasting & Climate Modeling

* **What It Is**: Simulations of the Earth’s atmosphere, oceans, and land processes.
* **Why It’s Intensive**: These models involve solving partial differential equations across 3D grids that can contain billions of cells. Tiny time steps are used for accuracy, leading to large computational workloads.
* **Examples**:
  * **Weather forecasting research, hurricane simulations, early warning systems.**

#### Computational Fluid Dynamics (CFD)

* **What It Is**: Numerical analysis and data structures to analyze fluid flows—key in engineering (aerospace, automotive).
* **Why It’s Intensive**: Accurate CFD often requires extremely fine meshes or grids to capture turbulence and boundary layers, resulting in massive numerical calculations.
* **Examples**:
  * **Aircraft Design**
  * **Vehicle Aerodynamics** (simulating airflow around cars for optimizing fuel economy).

#### Astrophysics & Cosmology

* **What It Is**: Simulating large-scale structures in the universe (e.g., galaxy formation), star evolution, black holes, and gravitational waves.
* **Why It’s Intensive**: Interactions among billions of particles or elements, along with complex physical laws (general relativity, plasma physics), make these simulations extremely heavy.
* **Examples**:
  * **Simulations of Galaxy Clusters** by universities and research institutes.
  * **Studying Black Hole Mergers**

#### Nuclear & Particle Physics

* **What It Is**: Modeling subatomic particle interactions, nuclear reactor cores, or accelerator experiments.
* **Why It’s Intensive**: Requires quantum-level physics, Monte Carlo methods, and/or large-scale iterative solvers.
* **Examples**:
  * **Simulating particle collisions**
  * **Fusion reactor simulations**

***

### 3. Visual Effects and 3D Rendering

#### Film & Animation Rendering

* **What It Is**: Creating photorealistic images and animations.
* **Why It’s Intensive**: Global illumination, ray tracing, and advanced physics-based lighting models require evaluating complex mathematical functions for each pixel. Frames can take hours each to render at high quality.
* **Examples**:
  * **Production** of feature films.
  * **Blender’s Cycles** (open-source) for CPU/GPU path tracing.

***

### 4. Protein Folding & Other Bioinformatics

#### Protein Structure Prediction

* **What It Is**: Determining how a protein’s amino acid chain folds into a 3D structure—crucial for understanding biological functions and designing drugs.
* **Why It’s Intensive**: The potential configuration space is astronomically large. Advanced methods (like AlphaFold) use deep learning models that require significant GPU resources.
* **Examples**:
  * **Predicting structures for nearly all known proteins.**
  * **Volunteer computing for protein folding research.**

#### Genome Sequencing & Assembly

* **What It Is**: Processing raw sequencing reads to reconstruct whole genomes (e.g., human, plant, bacterial).
* **Why It’s Intensive**: Datasets can easily reach terabytes in size. Algorithms like de novo assembly or alignment-based methods (e.g., Bowtie, BLAST) require large HPC clusters.
* **Examples**:
  * **Large-scale sequencing projects** (e.g., 1000 Genomes, Cancer Genomics).
  * **Metagenomic Studies** analyzing entire microbial communities.

#### Molecular Dynamics Simulations

* **What It Is**: Simulating the movement of atoms in molecules or complexes over time.
* **Why It’s Intensive**: Calculating forces and interactions at each step for millions of atoms demands CPU/GPU acceleration (e.g., GROMACS, NAMD).
* **Examples**:
  * **Drug Discovery** (predicting how small molecules bind to protein targets).
  * **Basic Biophysics Research** on membrane channels or virus capsids.

***

### 5. Big Data Analytics & Data Processing

#### Distributed Computing Frameworks

* **What It Is**: Systems like Apache Hadoop and Spark split massive datasets across many nodes for parallel processing.
* **Why It’s Intensive**: Operations like sorting, aggregating, or joining large tables can involve scanning petabytes of data.
* **Examples**:
  * **ETL Pipelines** for enterprise data lakes.
  * **Social Media Analytics** at companies like Twitter or LinkedIn, which process billions of events daily.

#### Graph Processing

* **What It Is**: Analyzing node-link structures to find patterns (e.g., community detection, shortest paths, graph embeddings).
* **Why It’s Intensive**: Graph algorithms can be complex (e.g., PageRank), with large real-world graphs (millions or billions of nodes/edges).
* **Examples**:
  * **Social Graph** for friend recommendations.

#### Real-Time Stream Processing

* **What It Is**: Handling data that arrives continuously (logs, sensor data, click streams) for immediate analytics and alerts.
* **Why It’s Intensive**: Requires fast ingestion, transformations, and real-time dashboards. Latency constraints necessitate efficient CPU/GPU usage.
* **Examples**:
  * **Financial Tick Data** in high-frequency trading.
  * **IoT Sensor Streams** in factories or connected devices.

***

### 6. Cryptography & Security

#### Encryption / Decryption at Scale

* **What It Is**: Securing large volumes of data in motion (TLS/SSL connections) and at rest (disk encryption).
* **Why It’s Intensive**: Bulk operations on huge data sets, but modern CPU instruction sets (AES-NI) and hardware accelerators help.
* **Examples**:
  * **VPN Gateways** handling encrypted connections for thousands of users.

#### Password Cracking & Security Auditing

* **What It Is**: Testing password strength by trying many possibilities (brute force) or using dictionary-based attacks.
* **Why It’s Intensive**: GPU-acceleration (e.g., via Hashcat) can test billions of hashes per second.
* **Examples**:
  * **Penetration Testing** for corporate security.
  * **Law Enforcement** accessing encrypted devices with court authorization.

***

### 7. Financial Modeling & Quantitative Analysis

#### Monte Carlo Simulations

* **What It Is**: Statistical simulations for risk assessment, derivative pricing (e.g., options, bonds), and portfolio optimization.
* **Why It’s Intensive**: Accurate results often require millions of iterations, each involving complex financial models.
* **Examples**:
  * **Derivative Pricing** of complex instruments (e.g., exotic options).
  * **Value at Risk (VaR)** calculations across large portfolios in investment banks.

#### High-Frequency Trading (HFT)

* **What It Is**: Automated trading strategies that react to market changes in microseconds or nanoseconds.
* **Why It’s Intensive**: The time factor is critical. Firms invest in specialized HPC clusters, FPGAs, or ASICs to reduce latency.
* **Examples**:
  * **Quant Funds** (Renaissance Technologies, Two Sigma) using large computing clusters.
  * **Market Making** requiring real-time price updates and predictive models.

***

### 8. Computer-Aided Design & Engineering (CAD/CAE)

#### Finite Element Analysis (FEA)

* **What It Is**: Breaking down complex structures into smaller elements to analyze stress, strain, heat transfer, etc.
* **Why It’s Intensive**: Large models with fine meshes require iterative solvers and matrix operations. Parallel processing across CPU cores or GPUs is often used.
* **Examples**:
  * **Automotive Crash Simulations** (ANSYS, LS-DYNA).
  * **Aerospace Structural Analysis**

#### Generative Design & Topology Optimization

* **What It Is**: Algorithms that iteratively suggest new designs based on performance goals (weight, strength, efficiency).
* **Why It’s Intensive**: Each iteration requires an analysis, which is repeated dozens or hundreds of times.
* **Examples**:
  * **Light weighting** automotive parts for better fuel efficiency.
  * **Architectural Design** for optimized building layouts (Autodesk’s Generative Design).

***

### 9. Video Encoding & Transcoding

#### High-Resolution Video (4K/8K)

* **What It Is**: Encoding large video files for streaming or storage using codecs like H.264, HEVC, VP9, AV1.
* **Why It’s Intensive**: Each frame undergoes complex compression algorithms. More pixels (4K/8K) = more data. Real-time or batch encoding at scale can stress CPU/GPU clusters.
* **Examples**:
  * **Media encoding** to optimize storage and delivery for various screen sizes.

***

### 10. Real-Time Simulation & Digital Twins

#### Smart City & Factory Simulations

* **What It Is**: Digital replicas of physical environments (e.g., a production line, traffic network) coupled with real-time sensor data.
* **Why It’s Intensive**: Simulating thousands of moving parts, IoT data streams, and complex event processing requires HPC-level resources.
* **Examples**:
  * **Digital twins for factory floors**, integrating robotics and sensor feedback.
  * **Urban Planning** tools simulating traffic flow, infrastructure loads, and environmental impact.

#### Automotive & Aerospace Digital Twins

* **What It Is**: Real-time virtual counterparts of vehicles or aircraft to test system updates, maintenance, and design changes.
* **Why It’s Intensive**: Must incorporate physics, AI-based control systems, and potentially large sensor datasets from the real-world counterpart.
* **Examples**:
  * Sports teams running simulations during races for strategy.


