AI agents are changing the amount of work that can happen inside a Git repository.
A developer might create a branch, make several changes, run tests, and push once or twice during a normal development session.
An agent can behave very differently.
It may modify files, run tests, checkpoint progress, create commits, push changes, inspect feedback, make another change, and repeat the process continuously.
That creates a new infrastructure problem:
Git repositories were designed around human development patterns. Agentic development can generate far more concurrent reads and writes.
GitHub is now rebuilding parts of its Git infrastructure around that reality while continuing to operate the existing platform. The goal is to support much higher repository activity without sacrificing reliability, consistency, governance, or the workflows developers already use.
The change is important because the bottleneck is no longer simply storing more Git repositories.
It is handling:
Developers
+
CI/CD
+
Code scanning
+
AI agents
|
v
Massive concurrent Git activity
Agentic Development Changes Git Workloads
Traditional Git usage has a relatively human-shaped workload.
A developer might:
Clone
|
v
Edit
|
v
Commit
|
v
Push
|
v
Pull Request
An agent can create a much tighter loop:
Inspect repository
|
v
Change code
|
v
Run tests
|
v
Commit/checkpoint
|
v
Push
|
v
Read feedback
|
v
Change code again
|
+--------> Repeat
That difference matters at scale.
GitHub reports that total Git activity more than doubled between September 2025 and August 2026, reaching more than 473 billion events per month. In September 2026, developers and agents created 7.38 billion commits, more than five times the volume a year earlier. Pushes grew from approximately 0.69 billion to 3.35 billion per month.
These numbers demonstrate why infrastructure designed around conventional developer behavior needs to evolve.
The Problem Is Not Just More Traffic
It is tempting to think that GitHub simply needs more servers.
That works reasonably well for read-heavy systems.
For example:
Client
|
v
Load Balancer
|
+--> Server A
+--> Server B
+--> Server C
+--> Server D
If all servers can read the same underlying data, adding capacity is relatively straightforward.
Git writes are different.
A push changes repository state.
That state must be:
stored durably
made consistent
visible to subsequent operations
available to CI
available to pull requests
available to APIs
protected against failures
A system cannot simply duplicate the write and hope every copy eventually agrees.
This makes write scalability much harder than read scalability.
Why Push Latency Matters More With Agents
For a human developer, a push taking an additional fraction of a second may be barely noticeable.
For an agent performing hundreds of operations, the same delay can become a meaningful part of the total execution time.
Consider an agent performing:
Action 1 -> Push
Action 2 -> Push
Action 3 -> Push
...
Action 100 -> Push
If every push waits for a relatively expensive commit path, the accumulated latency can become significant.
The infrastructure therefore needs to optimize not only total throughput but also the latency of individual writes.
GitHub describes this as a key challenge for agent-scale development: an agent's progress can be bounded by how quickly an individual push completes.
Write Throughput Is Becoming a Different Problem
GitHub reports that pushes increased roughly 4.9 times year over year, reaching about 3.35 billion pushes per month.
Imagine thousands of agents working against one large repository.
Each agent might have its own:
Branch
Commits
Pushes
Pull Requests
CI runs
Code scans
The repository becomes a shared coordination point.
The architecture therefore has to support:
Agent 1 ----\
Agent 2 -----\
Agent 3 ------> Repository
Agent 4 -----/
Agent 5 ----/
without allowing the repository's central state to become a serialization bottleneck.
Merge Operations Create Another Hotspot
Branches allow developers and agents to work independently.
Eventually, however, their changes need to converge.
For many organizations, that means a common default branch:
Agent A ----\
Agent B -----\
Agent C ------> main
Agent D -----/
Agent E ----/
The problem is that the default branch becomes a highly contested reference.
GitHub notes that pull request merges have grown to nearly four times their volume from a year earlier.
This creates a classic distributed-systems problem:
How much coordination is actually necessary?
If every operation requires coordination on the same resource, increasing the number of workers eventually stops helping.
Minimize Coordination
One of GitHub's architectural principles is to coordinate only where coordination is actually necessary.
For a Git push, the reference update is the critical consistency operation.
The underlying Git objects can often be processed independently.
Conceptually:
Push
|
+--> Store objects -----------+
|
+--> Validate objects --------+--> Parallel work
|
+--> Security scanning -------+
|
+--> Update reference --------> Coordination point
The idea is to keep the critical section small.
This is a familiar distributed-systems principle:
Do not serialize work that does not need to be serialized.
If object storage, validation, and other processing can happen concurrently, the system can reduce the amount of time that the central repository state remains a bottleneck.
Separating Storage From Compute
GitHub's existing infrastructure stores complete repository copies on local disks across multiple fileservers.
That approach provides useful properties:
Fast local reads
+
Durability
+
Redundancy
But it creates a trade-off.
If more copies are added to increase read capacity, those copies also participate in write operations.
That means:
More replicas
|
v
More read capacity
|
+
v
More write coordination
At extreme scale, the same replication mechanism that improves read availability can make writes more expensive.
GitHub's new architecture separates durable repository storage from the compute layer that serves Git requests.
The New Model
The architectural direction can be simplified as:
Git Requests
|
v
Compute Workers
/ | \
/ | \
Cache Cache Cache
\ | /
\ | /
\ | /
Durable Storage
The important distinction is that compute workers do not each need to be authoritative copies of the repository.
The durable storage layer becomes the source of truth.
Compute can then scale independently.
Why This Helps Read Scaling
Consider a repository receiving a large burst of CI activity.
A single push can trigger many downstream operations:
Push
|
+--> CI job 1
+--> CI job 2
+--> CI job 3
+--> Code scanning
+--> Test matrix
+--> Deployment
+--> Agent
+--> Pull request processing
GitHub says a single push can multiply into thousands of reads, while GitHub Actions alone ran 3.26 billion times in September 2026.
This is a fan-out problem.
A durable storage layer plus lightweight read workers can absorb that demand without requiring every additional read replica to become another participant in the write quorum.
The architecture becomes:
Push
|
v
Durable Storage
|
+----------+----------+
| | |
Worker A Worker B Worker C
| | |
CI reads Agent reads API reads
Read capacity can increase without proportionally increasing durable write coordination.
Maintenance Should Leave the Serving Path
Git repositories are not static files.
Git continuously creates and manages objects.
Repositories also require maintenance such as:
Object compaction
Garbage collection
Cleanup
Optimization
If maintenance runs on the same infrastructure responsible for serving live Git requests, a busy repository can experience contention.
The new architecture moves heavy maintenance work into separate workers operating directly against durable storage.
Conceptually:
Live Git traffic
|
v
Serving workers
|
v
Fast request path
Repository maintenance
|
v
Background workers
|
v
Durable storage
This is another example of separating latency-sensitive work from background work.
Why This Architecture Helps Failure Recovery
Consider a traditional replicated setup.
If a fileserver fails:
Fileserver failure
|
v
Reduced read capacity
+
Potential durability impact
Recovery may require rebuilding a full repository copy.
With storage and compute separated:
Compute worker fails
|
v
Replace worker
|
v
Read data from durable storage
|
v
Rebuild cache progressively
GitHub describes this model as making a compute-worker failure closer to a cache miss rather than a loss of an authoritative repository copy.
That is a significant architectural distinction.
Capacity Can Follow Demand
Traditional infrastructure often provisions enough compute for peak demand.
That can result in:
Normal traffic
|
v
Large idle capacity
A decoupled architecture allows compute workers to scale according to demand.
For example:
Normal day
|
v
10 workers
Release event
|
v
50 workers
Release complete
|
v
10 workers
The authoritative repository data does not need to be replicated every time compute capacity changes.
This makes the infrastructure more elastic.
Agent-Scale Development Is a Concurrency Problem
The deeper lesson is that AI agents do not simply create more requests.
They create more simultaneous software-development workflows.
A human team might have:
20 developers
10 CI jobs
An organization using thousands of agents can create:
Thousands of concurrent branches
Thousands of pushes
Thousands of test runs
Thousands of repository reads
The repository becomes an extremely busy shared state system.
That changes the engineering problem from:
How do we store Git repositories?
to:
How do we coordinate and serve repository state when thousands of independent workers are changing it concurrently?
That is a distributed-systems problem.
The Architecture Still Has to Preserve Git Semantics
Scaling Git infrastructure does not mean abandoning the semantics that developers depend on.
A push still needs to behave correctly.
A branch update must not accidentally overwrite another valid update.
A merge must respect repository history.
Pull requests must see consistent repository state.
The system therefore cannot simply eliminate coordination.
Instead:
Required coordination
|
v
Keep it
|
+
|
Unnecessary coordination
|
v
Remove it
This distinction is central to scaling write-heavy systems.
Governance Cannot Be an Afterthought
Agent-scale development also creates governance challenges.
More code changes mean more opportunities for:
Incorrect changes
Security issues
Accidental modifications
Unreviewed code
Policy violations
The infrastructure therefore has to preserve controls such as:
Branch protection
Required reviews
Audit logs
Repository permissions
Security scanning
Merge controls
GitHub explicitly states that its new infrastructure is intended to preserve the workflows and controls teams already depend on while supporting much higher activity levels.
That is important.
A faster Git platform that weakens repository governance would not solve the actual enterprise problem.
Humans Still Need Control
Agentic development changes who performs individual operations, but it does not eliminate ownership.
A healthy workflow remains:
Agent
|
v
Creates change
|
v
Tests
|
v
Pull Request
|
v
Human review
|
v
Approval
|
v
Merge
The infrastructure needs to support this workflow at much higher throughput.
The objective is not:
Let agents modify everything without supervision.
It is:
Let agents perform more work while people retain meaningful control over what enters the codebase.
Why Faster Git Infrastructure Matters to Developers
Infrastructure improvements can sound distant from everyday development.
They are not.
If repository operations become bottlenecks, developers eventually experience:
Slow pushes
Slow fetches
Delayed CI
Long merge queues
Slow branch operations
Agent waiting
Deployment delays
When those operations become faster and more predictable, the entire development loop improves.
Consider an agent working through:
Code
|
v
Test
|
v
Commit
|
v
Push
|
v
CI
|
v
Feedback
|
v
Code
Every infrastructure delay is multiplied by the number of iterations.
Reducing Git latency therefore improves more than Git itself.
It improves the complete software-development feedback loop.
Internal Results Show the Potential
GitHub reports that its new architecture has achieved up to 35 times higher write throughput in internal benchmarks, while read capacity can scale independently with demand.
That number should not be interpreted as a universal 35x improvement for every repository.
It is an internal benchmark demonstrating the potential of the architectural approach.
The more important result is the removal of a fundamental coupling:
Old model:
Read scaling
|
v
More durable replicas
|
v
More write overhead
New model:
Read scaling
|
v
More compute workers
Write durability
|
v
Durable storage
Separating those responsibilities gives each layer more freedom to scale independently.
What This Means for Enterprise Engineering
Enterprise teams adopting AI agents should pay attention to the same architectural principles.
Your Git platform may not operate at GitHub's scale, but the pattern can appear inside large organizations.
For example:
Repository
|
+--> Developers
+--> Coding agents
+--> CI
+--> Security scanners
+--> Dependency bots
+--> Release automation
All of those systems compete for repository resources.
The lesson is to identify shared bottlenecks before adding more automation.
If adding 1,000 agents produces 1,000 times more work against a single shared service, that service becomes the next bottleneck.
Agent scaling therefore requires infrastructure scaling.
What Engineering Teams Should Learn From This
There are several useful architectural lessons beyond GitHub itself.
Separate durability from compute
A durable data layer and a scalable compute layer do not necessarily need to be the same thing.
Minimize synchronization
Only coordinate operations that genuinely require a consistent shared state.
Move maintenance away from latency-sensitive paths
Background compaction, cleanup, and optimization should not unnecessarily compete with production requests.
Design for fan-out
One write can trigger thousands of downstream reads.
Treat agents as concurrent workers
An agent is not just another developer account. It can continuously generate work with very different timing characteristics.
Preserve governance
Scaling throughput without preserving review, audit, and security controls creates a different class of problems.
Common Mistakes
Assuming more replicas solve every scaling problem
Replication can increase read capacity but can also increase write coordination.
Measuring only average latency
Agents can perform many operations in tight loops. Tail latency can become important.
Ignoring merge contention
Thousands of independent branches eventually converge on shared references.
Running maintenance on the critical path
Compaction and cleanup can compete with live Git operations.
Treating agents as normal human users
Agent workloads can generate much higher operation frequency and concurrency.
Scaling automation without scaling infrastructure
More agents, CI jobs, and scanners increase pressure on shared systems.
Removing governance to increase speed
Fast changes without review and policy controls increase organizational risk.
Advantages and Disadvantages
Advantages
Higher write throughput: The redesigned architecture targets much greater concurrency for Git writes.
Independent read scaling: Read capacity can grow without requiring additional durable repository copies.
Better failure recovery: Compute failures can be handled more like cache failures.
Improved agent performance: Faster push operations reduce latency in agent development loops.
Better elasticity: Compute capacity can follow workload demand.
Reduced maintenance contention: Repository maintenance can operate independently from live request serving.
Disadvantages
Greater architectural complexity: Separating storage and compute introduces more distributed-system components.
Cache management becomes important: Read workers must efficiently retrieve and cache repository data.
Consistency remains difficult: Git references still require carefully controlled coordination.
Migration complexity: Rebuilding infrastructure underneath a running platform is significantly harder than designing a new system from scratch.
Operational observability matters more: More independent components create more failure modes that need to be monitored.
What Comes Next for Agent-Scale Git
The important part of GitHub's work is not simply replacing one storage system with another.
It represents a broader shift in how source-control infrastructure must operate when software development becomes increasingly automated.
The development loop is changing:
Traditional:
Developer
|
v
Edit
|
v
Commit
|
v
Push
Agentic:
Agent
|
+--> Inspect
+--> Modify
+--> Test
+--> Commit
+--> Push
+--> Review feedback
+--> Modify again
|
+------> Repeat
The infrastructure underneath that loop must handle significantly more concurrency without turning shared repository state into a bottleneck.
GitHub's approach is to separate concerns that were previously tightly coupled:
Durable repository data
+
Scalable compute
+
Independent read capacity
+
Background maintenance
+
Minimal coordination
That is a general distributed-systems pattern, not just a Git optimization.
Summary
GitHub is rebuilding its Git infrastructure because agentic software development changes the workload that source-control systems must support.
The challenge is not simply more repositories or more users.
It is the combination of:
More agents
+
More commits
+
More pushes
+
More CI
+
More repository reads
+
More concurrent branches
+
More merges
GitHub reports that Git activity has grown dramatically, with billions of commits and pushes occurring each month. The company is responding by separating durable repository storage from compute, scaling reads independently, minimizing coordination, and moving maintenance away from the live serving path.
For developers, the practical result should be a faster and more resilient Git foundation for increasingly automated development workflows.
For architects, the deeper lesson is more valuable.
Agent-scale development turns source control into a high-concurrency distributed system.
When thousands of automated workers continuously read, modify, test, and push code, the architecture must be designed around concurrency rather than simply scaled around traditional human workloads.
GitHub's infrastructure work is an example of that transition: preserve the Git workflows and governance developers already trust, while redesigning the underlying system so the development loop can operate at a much higher level of concurrency.
Join the conversation! Your thoughts help the community grow.