Training a large machine learning model and developing the code around it are usually treated as separate jobs.

A data scientist might use one environment for notebooks, another machine for experimentation, and a separate GPU cluster for training. Moving code, datasets, dependencies, and configuration between those environments adds friction to a workflow that already involves expensive infrastructure.

Amazon SageMaker HyperPod has been moving toward a different model.

HyperPod Spaces allow developers to run interactive development environments directly on HyperPod EKS clusters. With the latest SageMaker Studio integration, data scientists and ML engineers can create, configure, start, stop, and open those Spaces directly from the Studio interface instead of managing them primarily through the HyperPod CLI or kubectl. AWS announced the Studio integration on October 6, 2026.

The two environments available through this workflow are JupyterLab and Code Editor.

The interesting part is not simply getting another notebook or browser-based IDE. The development environment runs on the same HyperPod infrastructure that can be used for large-scale AI training and inference, which can reduce the separation between experimentation and production-scale compute.

What Is a SageMaker HyperPod Space?

A Space is essentially an interactive development environment running on a SageMaker HyperPod EKS cluster.

Developers can use it for:

Python development
Jupyter notebooks
Data preparation
Model experimentation
Training scripts
Debugging
Shell commands
Git workflows
Cluster interaction

The Space has its own compute, storage, and image configuration. AWS describes it as a self-contained environment for running JupyterLab or Code Editor directly on the HyperPod cluster.

The architecture looks roughly like this:

                    SageMaker Studio
                           |
                           v
                  HyperPod Cluster
                           |
             +-------------+-------------+
             |                           |
             v                           v
       JupyterLab Space           Code Editor Space
             |                           |
             +-------------+-------------+
                           |
                           v
                    HyperPod Compute
                           |
             +-------------+-------------+
             |                           |
             v                           v
       Training Workloads          Inference Workloads

Instead of developing on a laptop and then moving the code to a remote cluster, the developer can work closer to the compute where the actual ML workload runs.

What Changed in SageMaker Studio?

HyperPod Spaces were already available before the latest Studio integration.

The earlier workflow relied primarily on the HyperPod CLI or Kubernetes commands to create and manage Spaces. That approach is useful for infrastructure teams because it provides detailed control, but it puts more operational work on data scientists and ML engineers.

The new Studio experience adds an IDE and Notebooks tab to the HyperPod cluster details page.

From there, users can:

Create a Space
Configure compute
Configure storage
Select an image
Start a Space
Stop a Space
Open JupyterLab
Open Code Editor
Delete a Space
View Space status

The interface also exposes resource information such as GPU and vCPU allocation, storage, application type, and access type.

That makes Space management much closer to the experience developers expect from a managed development platform.

JupyterLab on HyperPod

JupyterLab is useful when the workflow is notebook-heavy.

A typical ML development session might look like:

Load Dataset
     ↓
Explore Data
     ↓
Transform Data
     ↓
Run Experiment
     ↓
Train Model
     ↓
Evaluate Results
     ↓
Adjust Parameters
     ↓
Run Again

Doing this directly on HyperPod means the notebook can operate closer to the GPU and storage resources used by the larger workload.

AWS describes the JupyterLab Space as providing Python notebooks, terminal access, file browsing, persistent storage, Spark-related kernels, and built-in development assistance.

The practical benefit is simple: the notebook becomes part of the same compute environment rather than just being a client sitting outside the ML infrastructure.

Code Editor for Developers Who Prefer an IDE

Not every ML engineer wants to build an application inside notebooks.

For larger projects, a traditional source-code workflow is often easier to maintain.

The HyperPod Code Editor Space provides a browser-based development environment with features such as:

Source-code editing
Syntax highlighting
IntelliSense
Integrated terminal
Git integration
Language extensions
Linters
Formatters
Cluster access

AWS describes it as a lightweight web-based IDE that can also access the cluster file system and mounted FSx volumes.

This makes a workflow such as the following possible:

Git Repository
     ↓
Code Editor
     ↓
Training Script
     ↓
HyperPod GPU
     ↓
Training Job
     ↓
Model Evaluation

The developer does not have to copy the project to a separate training machine just to run it.

You Can Also Use Your Local IDE

A browser-based IDE is not the only option.

HyperPod Spaces can also be accessed from a local IDE such as Visual Studio Code, Cursor, or Kiro through a remote connection. AWS recommends SSH over Systems Manager Session Manager for this workflow because it tunnels the connection through SSM without requiring inbound port 22 to be exposed on the cluster.

The model becomes:

Local VS Code
      |
      | SSH over SSM
      v
HyperPod Space
      |
      v
HyperPod Compute

This is useful for developers who prefer their local extensions, keyboard shortcuts, and editor configuration but still need their code to execute on remote GPU infrastructure.

There is an important distinction here.

The editor may be running locally, but the application code is running inside the HyperPod Space.

That can make a significant difference when the project depends on large datasets, specialized GPU libraries, or cluster-local resources.

Persistent Storage Separates Data From Compute

One of the useful characteristics of Spaces is that development data can persist independently from the compute environment.

AWS documents persistent EBS storage for Spaces, allowing users to stop and restart environments without losing work stored on the attached volume.

That gives the lifecycle a useful separation:

                 Space
                   |
        +----------+----------+
        |                     |
        v                     v
     Compute              Persistent Data
        |                     |
     Start/Stop           Keep Data

This matters for cost control.

A developer does not necessarily need to keep an expensive GPU-backed development environment running continuously.

A Space can be stopped when it is no longer needed while its persistent storage remains available.

GPU Allocation Can Be More Granular

GPU infrastructure is expensive, particularly when interactive development consumes an entire accelerator for relatively small workloads.

HyperPod Spaces support GPU allocation options that can be configured for the development environment. AWS also documents NVIDIA Multi-Instance GPU support for fractional GPU workloads on supported hardware.

For example, an organization might have:

Large GPU
    |
    +---- Space A
    |
    +---- Space B
    |
    +---- Training Workload

The exact allocation depends on the available hardware and cluster configuration, but the principle is important.

Not every developer experiment needs an entire accelerator.

Fractional allocation can make interactive development more practical when workloads are small enough to share GPU capacity.

HyperPod Task Governance Adds Another Layer

Shared GPU infrastructure introduces a familiar problem.

What happens when one developer consumes all the available capacity?

HyperPod Spaces can integrate with HyperPod Task Governance, allowing organizations to apply resource quotas and scheduling controls to interactive workloads. AWS exposes namespace-based configuration and governance when creating Spaces through Studio.

This can be structured around teams:

ML Platform
    |
    +---- Team A Namespace
    |        |
    |        +---- Developer 1
    |        +---- Developer 2
    |
    +---- Team B Namespace
             |
             +---- Developer 3
             +---- Developer 4

The platform team can then control how much compute each group is allowed to consume.

This becomes increasingly important as HyperPod moves from being a specialized training cluster toward shared AI development infrastructure.

The Same Environment Can Support Training and Development

The biggest architectural benefit is proximity.

Imagine a team training a large model on HyperPod.

Without Spaces:

Developer Laptop
       |
       v
Source Code
       |
       v
Build / Upload
       |
       v
Training Cluster
       |
       v
Collect Results
       |
       v
Developer Laptop

With Spaces:

HyperPod
   |
   +---- Development Space
   |
   +---- Training Workload
   |
   +---- Model Storage
   |
   +---- Shared File Systems

The developer can experiment directly where the compute exists.

AWS specifically describes HyperPod Spaces as a way to run interactive workloads alongside training and model deployment on the same infrastructure.

That can simplify workflows involving large datasets and shared file systems.

FSx and Shared Data

AI workloads frequently deal with datasets that are too large or inconvenient to repeatedly copy into individual development environments.

HyperPod Spaces can work with mounted file systems such as Amazon FSx. AWS documents access to mounted FSx volumes from Code Editor Spaces and also describes shared storage between development and training workflows.

A workflow might therefore look like:

                Shared FSx
                    |
        +-----------+-----------+
        |                       |
        v                       v
 Development Space        Training Workload
        |                       |
        +-----------+-----------+
                    |
                    v
              Model Artifacts

This can reduce unnecessary data movement and keep development and training closer to the same data environment.

It also means storage permissions become an important part of the architecture.

Developers should not automatically receive access to every dataset simply because they can launch a Space.

Startup Time Is Still a Consideration

There is an important trade-off with running Spaces on Kubernetes-backed infrastructure.

If the HyperPod cluster has scaled down and a developer creates a new Space, the environment may need to wait for compute capacity to become available.

AWS currently documents a cold-start delay of around five to seven minutes for Spaces on scale-to-zero clusters using Karpenter, with much of the delay coming from EC2 instance startup, Kubernetes node registration, and image pulling.

AWS also describes a node over-provisioning approach that can reduce Space startup to roughly 30–40 seconds by maintaining a pool of pre-warmed, image-cached nodes.

That creates a straightforward cost-versus-latency decision:

Scale to Zero
    |
    +---- Lower idle cost
    |
    +---- Longer startup

Pre-warmed Nodes
    |
    +---- Faster startup
    |
    +---- Some idle compute cost

For occasional experimentation, scale-to-zero may make more sense.

For teams repeatedly starting interactive sessions throughout the day, pre-warming may be worth the additional cost.

Stopping Spaces Helps Control Costs

A GPU-backed development environment does not need to run when nobody is using it.

The Studio integration makes it possible to stop Spaces directly from the management interface. AWS notes that stopping a Space releases compute resources while preserving data stored on the persistent EBS volume.

This makes the lifecycle more explicit:

Create
  ↓
Start
  ↓
Develop
  ↓
Stop
  ↓
Preserve Data
  ↓
Start Again

AWS also documents idle-shutdown capabilities for Spaces to help prevent unused interactive environments from consuming resources indefinitely.

For platform teams, this is an important operational control rather than just a convenience feature.

Security and Governance Still Matter

Moving development onto a shared GPU cluster creates new security considerations.

A platform team needs to think about:

Identity
IAM permissions
EKS access
Namespace boundaries
Storage permissions
Container images
Network access
Secrets
GPU quotas
Audit logs

For example, two teams may use the same HyperPod cluster but should not necessarily be able to access each other's development environments or datasets.

Namespaces provide one mechanism for organizing Spaces by team, while resource quotas and governance settings can control compute allocation.

The remote IDE workflow also deserves attention. AWS recommends SSM-based access because it avoids exposing inbound SSH ports on the cluster.

The important point is that an easier developer experience should not mean removing the platform controls around the infrastructure.

Common Mistakes

Treating a Space Like a Normal Laptop

A HyperPod Space is a development environment attached to shared infrastructure.

Developers still need to think about compute quotas, shared storage, cluster resources, and cloud costs.

Leaving GPU Spaces Running

An interactive GPU environment that is no longer being used can continue consuming expensive capacity.

Use stop controls or idle-shutdown policies where appropriate.

Giving Every Developer the Same Cluster Permissions

Developers may need access to their Space without needing administrative access to the underlying EKS cluster.

Keep developer and platform permissions separate.

Ignoring Storage Boundaries

Shared file systems make large datasets easier to access, but they also make access-control mistakes more consequential.

Apply data permissions independently of compute permissions.

Assuming Startup Is Instant

Scale-to-zero infrastructure can introduce several minutes of startup latency.

If developers need frequent interactive sessions, evaluate pre-warmed capacity.

Advantages and Disadvantages

Advantages

Development moves closer to AI compute. Developers can work directly on the same HyperPod infrastructure used for training and inference rather than repeatedly moving projects between environments.

JupyterLab and Code Editor support different workflows. Data scientists can use notebooks while software-oriented ML engineers can work in a more conventional IDE.

Local IDEs can still be used. Developers can connect VS Code and other supported local IDEs to the remote Space without exposing an inbound SSH port when using the SSM-based approach.

Compute can be managed independently from persistent data. Stopping a Space can release compute resources while preserving work stored on persistent storage.

GPU resources can be shared more efficiently. Fractional GPU allocation and governance features can help organizations make better use of expensive accelerator capacity.

Disadvantages

HyperPod is still specialized infrastructure. Teams need Kubernetes, AWS, networking, IAM, and ML platform knowledge to operate it effectively.

Cold starts can be significant. Scale-to-zero clusters can take several minutes to provide a usable Space.

The platform introduces additional governance requirements. Shared GPUs, datasets, namespaces, and persistent storage need clear access policies.

The best setup depends on workload patterns. Pre-warmed nodes can improve startup time but introduce additional infrastructure cost.

When HyperPod Spaces Make Sense

HyperPod Spaces are most useful for teams that already have meaningful workloads running on SageMaker HyperPod.

Typical examples include:

Foundation model development
Large-scale model training
GPU-heavy experimentation
Distributed ML workloads
AI research teams
Shared ML platforms
Data science teams using large datasets

They make less sense for a small ML project that runs comfortably on a local machine or a standard cloud development environment.

The value appears when the development workload itself needs access to the same specialized compute, storage, and cluster resources used by the larger AI system.

Summary

SageMaker HyperPod Spaces are turning HyperPod into more than a place where large training jobs run.

With the new SageMaker Studio integration, developers can create and manage JupyterLab and Code Editor environments directly from Studio, configure their compute and storage, start or stop Spaces, and open them in the browser without relying on the command line for everyday management.

The more important architectural change is that interactive development can happen alongside large-scale AI workloads on the same HyperPod infrastructure.

For data scientists, that means notebooks can operate closer to GPU compute and shared data. For ML engineers, Code Editor provides a more traditional development workflow. For platform teams, namespaces, quotas, persistent storage, GPU allocation, and governance provide mechanisms for operating these environments as a shared service.

There are still trade-offs. Scale-to-zero can introduce startup delays, GPU capacity is expensive, and shared infrastructure requires careful identity and data-access controls.

For organizations already investing in HyperPod, however, Spaces make the development side of the platform much more practical. The direction is clear: instead of treating experimentation, training, and inference as completely separate environments, AWS is bringing more of the AI development lifecycle onto the same underlying infrastructure.