Introduction
Rewriting a large production system in another programming language is rarely just a syntax conversion.
The engineering team has to understand the existing architecture, preserve behavior, maintain compatibility, improve performance where necessary, and avoid introducing new reliability problems during the transition.
GitHub's work on rewriting a substantial portion of its Copilot runtime in Rust provides an interesting example of this problem.
The project involved replacing a large amount of existing implementation with Rust while continuing to support a heavily used developer product.
The interesting part is not simply that Rust was selected. The more useful engineering lessons come from why a runtime becomes difficult to scale, how a rewrite can be structured, and what developers should consider before attempting a similar migration.
What Is the Copilot Runtime?
GitHub Copilot is an AI-assisted development product that must operate inside developer workflows.
A runtime supporting such a system has to coordinate several responsibilities:
Developer
|
v
Editor / IDE
|
v
Copilot Runtime
|
+---- Context Collection
+---- Request Management
+---- Network Communication
+---- Authentication
+---- Caching
+---- Model Interaction
+---- Response Handling
|
v
AI Services
The runtime therefore sits on a performance-sensitive path.
A delay or reliability problem in this layer can directly affect the developer experience.
Why Rewrite a Working System?
A rewrite is expensive.
If the existing implementation already works, the team needs a strong reason to replace it.
Common reasons for rewriting a runtime include:
Performance
Memory usage
Concurrency
Reliability
Startup time
Resource efficiency
Cross-platform requirements
Maintainability
Better control over low-level behavior
The decision should be based on measurable engineering constraints rather than language popularity.
A useful question is:
Which limitations of the existing implementation cannot be addressed economically without changing the underlying architecture or runtime model?
That is a much stronger starting point than:
Which programming language is better?
Why Rust Is Relevant to Runtime Engineering
Rust is designed around strong compile-time guarantees for memory safety without relying on a traditional garbage collector.
That makes it particularly interesting for systems where:
Low latency
+
High concurrency
+
Resource efficiency
+
Memory safety
are important.
Rust's ownership and borrowing model forces developers to make memory relationships explicit.
For runtime infrastructure, this can help prevent classes of memory-related errors before the program executes.
However, Rust does not automatically make an application faster.
Architecture, algorithms, I/O patterns, serialization, networking, and workload characteristics still determine real-world performance.
Rewriting 800,000 Lines Is an Engineering Problem
A rewrite of this scale cannot be treated as one giant migration.
A safer model is incremental:
Existing Runtime
|
v
Define Boundaries
|
v
Rewrite One Component
|
v
Compatibility Layer
|
v
Production Validation
|
v
Expand Migration
The old and new implementations may need to coexist for a period of time.
This allows the team to migrate functionality gradually instead of switching the entire system in one release.
Start With Architecture Boundaries
Before rewriting code, identify logical components.
For example:
Copilot Runtime
|
+---- Configuration
|
+---- Authentication
|
+---- Context
|
+---- Networking
|
+---- Caching
|
+---- Request Pipeline
|
+---- Response Processing
Each boundary becomes a potential migration unit.
The goal is to avoid creating a situation where every rewritten component depends directly on every remaining component.
Strong boundaries make migration possible.
Preserve Behavior Before Changing It
A common rewrite mistake is attempting to modernize architecture and change behavior simultaneously.
Suppose the original runtime performs:
Request
|
+---- Collect Context
|
+---- Send Request
|
+---- Process Response
|
v
Display Result
The first Rust implementation should ideally preserve this behavior.
Do not simultaneously introduce:
New caching semantics
New authentication rules
New protocol behavior
New request scheduling
New error semantics
unless those changes are independently required.
Behavioral compatibility makes debugging dramatically easier.
Define a Compatibility Contract
A rewrite needs a contract between the old and new implementations.
The contract can include:
Inputs
Outputs
Error conditions
Timeouts
Retry behavior
Authentication
Protocol formats
Logging
Telemetry
For example:
Input
|
v
Runtime API
|
+---- Success -> Expected Response
|
+---- Timeout -> Expected Timeout
|
+---- Failure -> Expected Error
This makes the new implementation testable against known behavior.
Use Differential Testing
One powerful migration technique is differential testing.
The same input is sent through both implementations:
Request
|
+-------+-------+
| |
v v
Old Runtime Rust Runtime
| |
v v
Result A Result B
| |
+-------+-------+
|
v
Compare
The results can then be compared automatically.
This approach is useful when rewriting code where exact behavior matters.
It can reveal:
Output differences
Error differences
Serialization differences
Timeout differences
Edge-case regressions
Rust Ownership Changes How Code Is Designed
A developer moving from a garbage-collected language to Rust encounters a different programming model.
Consider a simplified data structure:
struct RequestContext {
repository: String,
file_path: String,
content: String,
}
A function can borrow the structure rather than copying it:
fn build_prompt(context: &RequestContext) -> String {
format!(
"Repository: {}\nFile: {}",
context.repository,
context.file_path
)
}
The function does not take ownership of context.
This distinction matters when runtime code processes large amounts of context.
Developers need to think carefully about:
Ownership
Borrowing
Lifetimes
Allocation
Cloning
Shared state
Concurrency
Avoid Blindly Cloning Data
One of the easiest ways to undermine a Rust performance effort is excessive cloning.
For example:
let copy = context.content.clone();
Sometimes a clone is appropriate.
But if large strings, buffers, or context objects are repeatedly copied, memory usage can increase substantially.
Instead, consider whether a borrowed reference is sufficient:
fn process(content: &str) {
// Process without taking ownership.
}
The compiler's ownership rules encourage developers to make this decision explicitly.
Concurrency and Runtime Workloads
AI developer tooling can perform multiple operations concurrently.
For example:
Request
|
+---- Read Files
|
+---- Inspect Repository
|
+---- Gather Editor Context
|
+---- Prepare Network Request
|
+---- Update Cache
A runtime needs to coordinate these operations efficiently.
Rust's asynchronous ecosystem provides mechanisms for concurrent I/O workloads.
A simplified example might use asynchronous tasks:
async fn collect_context() -> Result<Context, Error> {
let files = load_files().await?;
let repository = load_repository_info().await?;
Ok(Context {
files,
repository,
})
}
Production systems require much more sophisticated scheduling, cancellation, error handling, and resource management.
The important design point is that concurrency should be explicit and controlled.
Cancellation Matters
AI requests can become unnecessary.
A developer may:
Type Code
|
v
Request Starts
|
v
Developer Changes Code
|
v
Previous Request No Longer Relevant
Continuing expensive work in the background wastes resources.
A runtime should therefore support cancellation where appropriate.
Conceptually:
async fn request(
cancellation: CancellationToken
) -> Result<Response, Error> {
cancellation.cancelled().await;
// Stop work when cancellation is requested.
}
The exact implementation depends on the asynchronous framework and architecture.
Cancellation must propagate through the entire operation instead of stopping only the outer function.
Memory Efficiency
Runtime systems often process large amounts of transient data.
Examples include:
Source files
Repository metadata
Editor buffers
Model responses
JSON payloads
Network buffers
A rewrite should measure allocation behavior instead of assuming the new language solves memory problems automatically.
Useful measurements include:
Peak Memory
Allocation Rate
Request Memory
Cache Size
Resident Memory
Startup Memory
These measurements should be collected under realistic workloads.
Startup Time Matters
Developer tools often run locally and interact with an IDE.
Startup latency can therefore affect every developer.
A runtime that takes several seconds to initialize may become noticeable when:
Opening an IDE
Opening a repository
Switching projects
Restarting after an update
Recovering from an error
A smaller runtime footprint can help, but startup performance should be measured end to end.
Networking Is Part of Runtime Performance
An AI coding runtime spends significant time communicating with remote services.
A useful latency model is:
Total Latency
=
Local Processing
+
Context Preparation
+
Network Latency
+
Server Processing
+
Response Processing
Optimizing only the local runtime will not solve a slow remote service.
The engineering team therefore needs measurements for each stage.
For example:
Stage | Measurement |
|---|---|
Context collection | Measure |
Serialization | Measure |
Network request | Measure |
Server response | Measure |
Response parsing | Measure |
UI delivery | Measure |
This helps prevent optimizing the wrong bottleneck.
Cross-Platform Runtime Design
Developer tools run across multiple operating systems.
A runtime may need to support:
Windows
macOS
Linux
Cross-platform code needs careful treatment of:
File paths
Processes
Environment variables
File watchers
Networking
Certificates
Permissions
Native dependencies
A rewrite should include platform-specific testing rather than assuming that successful compilation means equivalent behavior.
Observability During a Rewrite
A large migration needs strong telemetry.
Track both implementations while they coexist.
Useful signals include:
Request success rate
Error rate
Latency
Startup time
Memory consumption
CPU utilization
Cancellation rate
Network failures
Tool failures
A migration dashboard might look conceptually like:
Runtime Migration
|
+---- Old Runtime
| |
| +---- Error Rate
| +---- Latency
|
+---- Rust Runtime
|
+---- Error Rate
+---- Latency
This allows the team to compare implementations using production evidence.
Gradual Rollout
Do not move every user to the rewritten runtime immediately.
A safer approach is:
Internal Users
|
v
Small Percentage
|
v
Larger Percentage
|
v
Broad Rollout
At every stage, monitor:
Errors
Performance
Crash reports
User-impacting regressions
Resource usage
If problems appear, the rollout can be paused or reversed.
What Developers Can Learn From the Rewrite
The most important lesson is not that Rust should replace every runtime.
It is that language choice should follow system requirements.
For a performance-sensitive runtime, the engineering team may value:
Predictable resource usage
Memory safety
Concurrency
Low-level control
Small runtime footprint
For other systems, developer productivity, ecosystem maturity, or rapid iteration may dominate the decision.
There is no universal replacement language.
Common Mistakes in Large Rewrites
Rewriting Everything at Once
A complete cutover creates a huge failure surface.
Changing Architecture and Language Simultaneously
Too many variables make regressions difficult to diagnose.
Measuring Only Synthetic Benchmarks
A benchmark may not represent actual developer workloads.
Ignoring Memory Behavior
CPU performance alone does not describe runtime efficiency.
Underestimating Integration Work
The surrounding system can be harder to migrate than the core implementation.
Removing the Old Implementation Too Early
The old system may still be necessary as a behavioral reference during migration.
Treating Compiler Success as Migration Success
A program that compiles can still have different runtime behavior.
Best Practices for Large Runtime Rewrites
Define clear component boundaries.
Establish behavioral contracts before rewriting.
Migrate incrementally.
Use differential testing where possible.
Measure real workloads.
Track memory as well as CPU performance.
Design cancellation into asynchronous operations.
Test every supported operating system.
Maintain strong observability during migration.
Use staged production rollouts.
Keep rollback mechanisms available.
Avoid unnecessary data copying.
Separate language migration from unrelated feature changes.
Document compatibility decisions.
Remove legacy code only after the replacement has demonstrated stability.
Advantages of Rust for Runtime Components
Rust can provide several properties that are valuable for runtime infrastructure:
Compile-time memory-safety guarantees
No traditional garbage collector
Explicit ownership
Strong concurrency safety
Fine-grained control over resource usage
Good support for asynchronous systems
Native compilation
These properties can make Rust a strong candidate for certain runtime workloads.
Limitations and Tradeoffs
Rust also introduces costs.
Developers must learn:
Ownership
Borrowing
Lifetimes
Trait-based abstractions
Async runtime design
Rust-specific tooling
Migration itself is expensive.
An organization also needs:
Rust expertise
Build infrastructure
Debugging capability
Cross-platform testing
Long-term maintenance plans
A rewrite should therefore have measurable engineering objectives.
When Does a Rewrite Make Sense?
A large rewrite deserves serious consideration when the existing implementation has measurable limitations that justify the migration cost.
Potential signals include:
Performance bottleneck
+
Memory pressure
+
Concurrency limitations
+
Maintenance cost
+
Clear migration boundary
A rewrite is harder to justify when the primary motivation is simply that another programming language is currently popular.
A Practical Migration Checklist
Before starting a large runtime rewrite:
Identify the measurable problems.
Define success metrics.
Map the existing architecture.
Identify migration boundaries.
Define compatibility contracts.
Build automated regression tests.
Establish differential testing where possible.
Select a small component for the first migration.
Measure the new implementation.
Introduce staged rollout infrastructure.
Monitor production behavior.
Expand migration gradually.
Maintain a rollback strategy.
Remove legacy components only after sufficient validation.
Summary
GitHub's large-scale work on its Copilot runtime illustrates the engineering complexity behind replacing a mature production implementation.
The difficult part of a rewrite is not converting source code from one language to another. It is preserving behavior while improving the characteristics that motivated the rewrite in the first place.
Rust can provide useful properties for performance-sensitive runtime components, particularly around memory safety, concurrency, and resource control. But those benefits only matter when they address measurable system requirements.
For developers considering a similar migration, the strongest approach is incremental: define boundaries, preserve behavior, measure real workloads, compare old and new implementations, roll out gradually, and keep a reliable rollback path.
A successful rewrite is ultimately less about the number of lines converted and more about whether the new implementation delivers measurable improvements without compromising the product it supports.

Join the conversation! Your thoughts help the community grow.