Linux control groups, better known as cgroups, are one of the mechanisms Kubernetes uses to control how much CPU, memory, and other resources containers can consume.
For years, Kubernetes environments commonly ran on cgroup v1.
That model is now being phased out.
cgroup v2 is the newer Linux resource-management interface, has been stable in Kubernetes since 1.25, and Kubernetes has deprecated cgroup v1 since 1.35. The transition matters because cgroups are not just an implementation detail. They affect resource limits, quality of service, monitoring, runtime behavior, and how applications interpret the resources available inside a container.
The migration can look deceptively simple:
cgroup v1
|
v
cgroup v2But production Kubernetes clusters have several layers that depend on cgroups:
Linux Kernel
|
v
cgroup v2
|
+--> systemd
|
+--> container runtime
|
+--> kubelet
|
v
Kubernetes Pods
|
v
ApplicationsFor teams operating Kubernetes clusters, now is a good time to determine whether workloads, monitoring tools, runtimes, and node-level automation are ready for cgroup v2.
What Are Linux cgroups?
A cgroup is a Linux kernel mechanism for organizing processes and controlling their resource consumption.
Kubernetes uses cgroups to enforce resources specified in Pod and container configurations.
For example:
resources:
requests:
cpu: "250m"
memory: "256Mi"
limits:
cpu: "1"
memory: "512Mi"These values do not magically constrain the process.
The Kubernetes control plane and kubelet eventually translate resource configuration into controls implemented through the Linux operating system and container runtime.
Conceptually:
Pod Spec
|
v
CPU / Memory Requests & Limits
|
v
Kubelet
|
v
Container Runtime
|
v
Linux cgroups
|
v
Kernel Resource ControlThat is why changing the underlying cgroup model can affect application behavior even when the Kubernetes YAML remains unchanged.
cgroup v1 vs cgroup v2
The fundamental difference is architectural.
cgroup v1 uses multiple resource hierarchies. Different controllers can be organized through separate hierarchies.
cgroup v2 provides a single unified hierarchy.
A simplified comparison looks like this:
Area | cgroup v1 | cgroup v2 |
|---|---|---|
Hierarchy | Multiple hierarchies | Unified hierarchy |
Resource model | Separate controller model | Unified interface |
Memory accounting | Older model | More comprehensive accounting |
Resource isolation | Mature but fragmented | Improved unified model |
PSI | Not a cgroup-v1 feature | Available with v2 |
Kubernetes direction | Deprecated | Strategic default |
The unified design makes resource management more consistent and provides newer capabilities such as Pressure Stall Information and improved resource accounting.
Why Kubernetes Is Moving to cgroup v2
Kubernetes increasingly depends on capabilities that are better supported by cgroup v2.
One example is MemoryQoS, which uses cgroup v2 primitives to improve memory quality-of-service behavior.
cgroup v2 also provides improved accounting for different categories of memory and resource activity.
That gives Kubernetes a better foundation for controlling workloads under resource pressure.
The broader architecture becomes:
Kubernetes
|
+---------+---------+
| |
CPU control Memory control
| |
+---------+---------+
|
cgroup v2
|
Linux KernelInstead of maintaining separate mechanisms for different resource controllers, the system works through a unified hierarchy.
cgroup v2 Is Already Stable
This is not an experimental Kubernetes feature.
Kubernetes has supported cgroup v2 as a stable capability since version 1.25. Kubernetes documentation now describes cgroup v1 as deprecated beginning with Kubernetes 1.35.
That changes the migration conversation.
Teams should no longer think:
Should we experiment with cgroup v2?
The better question is:
What dependencies in our environment still assume cgroup v1?
That includes applications, monitoring agents, security tools, container runtimes, node-level scripts, and custom operational tooling.
The cgroup Driver Is a Separate Concept
One common source of confusion is mixing up:
cgroup versionwith:
cgroup driverThey are related but different.
The cgroup version describes the Linux kernel interface:
v1
v2The cgroup driver describes how components such as kubelet and the container runtime manage cgroups.
Common drivers include:
cgroupfs
systemdFor cgroup v2, Kubernetes recommends using the systemd cgroup driver rather than cgroupfs.
The architecture should therefore be consistent:
systemd
|
+--> kubelet
|
+--> container runtime
|
v
cgroup v2The kubelet and container runtime should use compatible cgroup-driver configuration.
Why systemd Matters
Modern Linux distributions commonly use systemd as their init system.
systemd already manages processes through cgroups.
If Kubernetes independently manages cgroups through cgroupfs, the machine can end up with two competing views of resource management.
That can become problematic under resource pressure.
Kubernetes recommends using systemd as the cgroup driver when systemd is the init system.
A kubelet configuration can look like:
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
cgroupDriver: systemdThe important point is not simply setting this value.
The container runtime must use the compatible driver as well.
Container Runtime Compatibility Matters
The cgroup stack includes the container runtime.
A typical Kubernetes node looks like:
Kubelet
|
v
CRI
|
v
containerd / CRI-O
|
v
Linux cgroup v2If the kubelet and runtime disagree about cgroup management, the node can behave unpredictably.
Modern Kubernetes versions are improving automatic cgroup-driver detection through the CRI.
Kubernetes documents automatic detection through the RuntimeConfig CRI RPC, with support from modern runtimes such as containerd 2.0+ and CRI-O 1.28+.
This reduces configuration drift, but older environments may still require explicit configuration.
How to Check Your Linux Node
The simplest way to determine whether a Linux system is using cgroup v2 is:
stat -fc %T /sys/fs/cgroup/For cgroup v2, the output is:
cgroup2fsFor cgroup v1, it typically reports:
tmpfsKubernetes documents this as a standard way to identify the active cgroup version.
On a node where you have appropriate access, you can also inspect:
mount | grep cgroupand:
cat /proc/cgroupsHowever, /proc/cgroups alone is not the best way to determine the active cgroup hierarchy on a modern system. The filesystem check is more direct.
What Changes for Applications?
For many containerized applications, nothing needs to change in application code.
That is one of the important characteristics of the migration.
Your Kubernetes deployment can continue to specify:
resources:
requests:
cpu: "500m"
memory: "512Mi"
limits:
cpu: "1"
memory: "1Gi"The application does not normally need to know whether the host is using cgroup v1 or v2.
The problem appears when software explicitly reads the cgroup filesystem.
For example:
Application
|
v
/sys/fs/cgroup/...An application or monitoring tool that assumes the old cgroup v1 directory layout can fail after migration.
GKE explicitly warns that workloads and third-party tools reading the cgroup filesystem need to be compatible with the cgroup v2 API.
Runtime Resource Detection Can Change
Applications sometimes inspect container limits to determine how much memory or CPU they should use.
This is particularly important for runtimes such as:
JVM
Node.js
.NET
Go
Python extensions
Native applicationsThe application might effectively ask:
How much memory is available to me?
If the runtime incorrectly reads the host's memory instead of the container's cgroup limit, it can make bad decisions.
For example:
Host memory: 64 GB
Pod memory limit: 2 GBIf the runtime thinks it has 64 GB available, it might allocate significantly more memory than the container can actually use.
That can lead to:
Memory pressure
|
v
OOM termination
|
v
Pod restartNode.js Is a Good Example
Kubernetes documentation specifically notes that Node.js versions starting with 20.3.0 can detect cgroup v2 memory limits through libuv, while the Node.js 18 line does not reliably detect them. Applications on affected versions may therefore size their heap incorrectly unless the heap size is explicitly constrained.
This creates a practical migration rule:
cgroup v2 migration
|
v
Check runtime resource detection
|
v
Check memory limits
|
v
Check OOM behaviorFor Node.js applications that still run older runtime versions, this is especially important.
Java Applications Also Need Validation
Java applications have historically performed container resource detection based on Linux cgroup information.
Modern JDK releases have strong cgroup v2 support, but teams running older JDK versions should verify compatibility before moving nodes.
GKE recommends Java versions that fully support cgroup v2, including JDK 8u372, JDK 11.0.16 or later, or JDK 15 and later.
The general principle applies beyond Java:
Do not assume that an application understands container limits simply because it runs inside a container.
Verify it.
Monitoring Agents Are a Major Migration Risk
Monitoring systems are often more tightly coupled to cgroups than application code.
For example:
Node exporter
Metrics agent
APM agent
Container monitoring
Security agent
Custom daemonmay read:
/sys/fs/cgroup/directly.
A v1-aware tool may expect paths such as:
/sys/fs/cgroup/memory/...while cgroup v2 exposes a unified hierarchy.
A migration can therefore produce a particularly dangerous failure:
Application is healthy
|
v
Monitoring becomes inaccurateThat is worse than an obvious application crash because operators may not notice the problem until a real incident occurs.
Before migration, identify every node-level tool that reads cgroup information.
Security Tools Need Testing Too
Security software can also interact with cgroups.
For example:
Runtime security
Process monitoring
Container isolation
Admission agents
Host intrusion detection
eBPF-based toolingSome tools may use cgroup identifiers to associate processes with containers.
If the tool expects a v1 layout, it may stop identifying workloads correctly after the migration.
GKE explicitly calls out monitoring and third-party tools as dependencies that should be verified for cgroup v2 compatibility.
Kubernetes Resource Requests Still Matter
Moving to cgroup v2 does not eliminate the importance of Kubernetes resource requests and limits.
For example:
resources:
requests:
cpu: "250m"
memory: "256Mi"
limits:
cpu: "1"
memory: "512Mi"The scheduler still uses requests to make placement decisions.
The runtime and kernel still enforce resource controls.
A useful mental model is:
Requests
|
v
Scheduling
Limits
|
v
Runtime enforcement
cgroup v2
|
v
Kernel resource controlcgroup v2 changes the underlying Linux resource-management model, not the basic purpose of Kubernetes resource requests and limits.
Memory Pressure Deserves Special Attention
Memory is usually where resource-management migrations become most visible.
Consider a Pod with:
resources:
limits:
memory: "512Mi"Under load, the application approaches its limit.
The system must decide how memory pressure is handled.
With cgroup v2, Kubernetes can use newer kernel mechanisms for improved resource management and isolation.
That does not mean memory limits become unlimited or that OOM kills disappear.
If an application genuinely requires more than:
512 MiBit still needs a higher memory limit or a more memory-efficient implementation.
The migration should therefore include memory-pressure testing.
Test CPU Behavior Too
CPU semantics differ internally between cgroup versions.
cgroup v1 uses CPU shares, while cgroup v2 uses CPU weight.
Kubernetes has continued improving how CPU requests are translated between these models, including work on the conversion from cgroup v1 CPU shares to cgroup v2 CPU weight.
For most applications, the migration should be transparent.
For CPU-sensitive workloads, however, benchmark:
CPU-bound workers
Batch jobs
Latency-sensitive APIs
High-density nodesespecially when multiple workloads compete for CPU.
Pressure Stall Information
One of the useful capabilities associated with cgroup v2 is Pressure Stall Information, or PSI.
PSI measures how much time workloads spend stalled because resources such as CPU, memory, or I/O are unavailable.
Conceptually:
Application wants CPU
|
v
CPU unavailable
|
v
Process waits
|
v
Resource pressureTraditional utilization tells you how much of a resource is being used.
Pressure information can provide another view:
How much time are workloads losing because the resource is unavailable?
This can be valuable for diagnosing node-level contention and understanding why an application is slow even when simple utilization metrics do not tell the whole story.
GKE Has Its Own Migration Timeline
For teams using GKE, the migration is particularly important.
Google's current GKE transition has several milestones:
GKE < 1.26
|
v
cgroup v1 default for nodes
GKE 1.26+
|
v
cgroup v2 default for new nodes
GKE 1.31
|
v
cgroup v1 deprecated
GKE 1.33+
|
v
Automatic migration of cgroup v1 clusters
GKE 1.35
|
v
cgroup v1 removedGoogle documents that new nodes on GKE 1.26 and later use cgroup v2 by default, while existing cgroup v1 node pools do not automatically switch merely because they are upgraded to 1.26. GKE begins migrating remaining cgroup v1 clusters starting with version 1.33 and removes cgroup v1 support in 1.35.
This means a cluster can be running a relatively modern Kubernetes version while some existing nodes still use the old cgroup mode.
Check GKE Before Assuming Anything
For GKE, determine the actual cgroup mode rather than inferring it from the cluster version.
For example, GKE provides commands to inspect the effective cgroup mode of cluster resources.
The important distinction is:
Cluster version
!=
Actual node cgroup modeA cluster can contain older node pools whose configuration was established before cgroup v2 became the default.
GKE documents separate behavior for Autopilot, manually managed Standard node pools, and node auto-provisioning.
GKE Node Pool Migration
For a Standard cluster, a node pool can be configured to use cgroup v2 through node system configuration.
Conceptually:
linuxConfig:
cgroupMode: 'CGROUP_MODE_V2'GKE then recreates nodes using the selected configuration.
That node recreation is important.
Do not think of this as:
Change setting
|
v
Same node continues runningThe practical migration is closer to:
Existing node
|
v
Drain / replace
|
v
New node
|
v
cgroup v2
|
v
Pods rescheduledTherefore, Pod disruption budgets, capacity, topology constraints, and workload availability all matter.
Do Not Change cgroup Drivers on Live Nodes Casually
Changing the cgroup driver of an already-running node is a sensitive operation.
Kubernetes warns that changing the driver can cause problems recreating existing Pod sandboxes, and simply restarting the kubelet may not resolve those problems. Replacing or reinstalling nodes through automation is the safer approach when possible.
A safer migration pattern is:
Create compatible node pool
|
v
Verify cgroup v2
|
v
Schedule workloads
|
v
Drain old nodes
|
v
Delete old node poolThis is more controlled than modifying the foundation underneath running workloads.
A Practical Migration Checklist
Before migration, inventory:
[ ] Kubernetes version
[ ] Linux distribution
[ ] Kernel version
[ ] Container runtime
[ ] cgroup version
[ ] cgroup driver
[ ] Monitoring agents
[ ] Security agents
[ ] APM agents
[ ] Node-level scripts
[ ] Runtime versions
[ ] Native applicationsThen test:
[ ] CPU limits
[ ] CPU requests
[ ] Memory limits
[ ] Memory pressure
[ ] OOM behavior
[ ] Autoscaling
[ ] Monitoring
[ ] Logging
[ ] Security tooling
[ ] Application startup
[ ] Graceful shutdownFinally:
[ ] Create replacement nodes
[ ] Validate workload behavior
[ ] Drain old nodes
[ ] Monitor production
[ ] Remove legacy configurationCommon Mistakes
Checking only the Kubernetes version
A newer Kubernetes version does not necessarily mean every existing node has already migrated.
Assuming applications cannot be affected
Applications that inspect cgroup files or container limits can behave differently.
Forgetting monitoring agents
A monitoring agent can fail silently while applications continue running.
Changing cgroup drivers in place
This can create Pod sandbox and runtime problems.
Migrating every node simultaneously
A rolling node replacement strategy is safer.
Ignoring runtime versions
Old Node.js, Java, or other runtimes may interpret container limits incorrectly.
Testing only normal workloads
Test memory pressure and CPU contention as well.
Ignoring third-party software
APM, security, and observability agents may have their own cgroup assumptions.
Advantages and Disadvantages
Advantages
Unified resource model: cgroup v2 provides a single hierarchy for resource management.
Improved isolation: The newer interface provides stronger and more consistent resource-management primitives.
Better memory accounting: cgroup v2 improves accounting across different types of memory activity.
Pressure information: PSI provides another way to understand resource contention.
Future Kubernetes alignment: cgroup v2 is the direction Kubernetes is taking as cgroup v1 is deprecated.
Newer resource-management capabilities: Features such as MemoryQoS rely on cgroup v2 primitives.
Disadvantages
Migration complexity: Existing clusters may contain nodes and tools designed around cgroup v1.
Tool compatibility: Monitoring, security, and APM software may directly inspect cgroup information.
Runtime compatibility: Older language runtimes may incorrectly detect container resource limits.
Operational risk: Changing node-level configuration can require node replacement.
Debugging complexity: cgroup v2 changes paths and resource-management semantics that operators may be accustomed to.
How C# and .NET Teams Should Prepare
For .NET applications running on Kubernetes, the migration is usually transparent at the application-code level.
You should still validate:
.NET runtime
|
+--> CPU limits
+--> Memory limits
+--> GC behavior
+--> Thread pool
+--> Container startup
+--> OOM behaviorRun the application under realistic resource limits.
For example:
resources:
requests:
cpu: "500m"
memory: "512Mi"
limits:
cpu: "1"
memory: "1Gi"Then test:
Normal traffic
High CPU
High allocation rate
Memory pressure
Pod restart
Horizontal scaling
Node replacementThe goal is to verify that the runtime's behavior matches the Kubernetes resource configuration.
A Better Migration Strategy
The safest strategy is not to perform a giant cluster-wide switch.
Use progressive migration:
Existing cluster
|
v
Identify cgroup v1 nodes
|
v
Create cgroup v2 node pool
|
v
Run representative workloads
|
v
Validate monitoring and security
|
v
Load test
|
v
Move production workloads
|
v
Drain cgroup v1 nodes
|
v
Remove legacy nodesThis approach provides a rollback point.
If a workload behaves unexpectedly, move it back to the known-good environment while investigating.
Summary
Kubernetes is moving decisively toward cgroup v2 as the standard Linux resource-management model.
The important change is not the filesystem layout alone.
cgroup v2 affects the underlying mechanism used to enforce CPU and memory resources, provides a unified hierarchy, enables newer resource-management capabilities, and gives Kubernetes a stronger foundation for features such as MemoryQoS.
For GKE users, the transition is already well underway. Newer GKE node pools use cgroup v2 by default, GKE is migrating remaining cgroup v1 environments, and cgroup v1 support is scheduled for removal with GKE 1.35.
For application teams, the safest approach is to focus on compatibility rather than the cgroup implementation itself.
Check:
Runtime
Monitoring
Security
APM
Container runtime
Node configuration
Resource limitsThen migrate nodes progressively and test workloads under real CPU and memory pressure.
Most well-behaved applications should continue running without application-code changes.
The workloads that require attention are the ones that make assumptions about the Linux cgroup filesystem, depend on old runtime resource detection, or use node-level tooling that has not yet been updated for cgroup v2.
The practical goal is simple:
Do not wait for cgroup v1 removal to discover which parts of your Kubernetes platform still depend on it.
Join the conversation! Your thoughts help the community grow.