Why Parallel Code Loses Updates: lock, Interlocked, and Thread-Local State, Measured

Concurrency advice is usually delivered as rules: "use lock", "prefer Interlocked", "avoid shared state". Rules are easier to remember when you have watched them fail. So I built a deliberately small experiment, incrementing one counter 100,000 times with Parallel.For, and measured what happens with an unsafe version and four correct ones. The results included one expected finding and two I did not expect.

Everything below marked as measured comes from my own runs on one laptop with .NET 8. Anything I did not measure is labeled as such.

1. The Test Bed

csharp

using System.Diagnostics;

const int N = 100_000;

void Time(string name, Func<long> work)
{
    var sw = Stopwatch.StartNew();
    long result = work();
    sw.Stop();
    Console.WriteLine($"{name,-22} {result,8}   {sw.Elapsed.TotalMilliseconds,8:F3} ms");
}

Parallel.For(0, N, i => { });   // warm-up, not measured

for (int round = 1; round <= 3; round++)
{
    Console.WriteLine($"--- round {round} ---");

    Time("Sequential", () => { long c = 0; for (int i = 0; i < N; i++) c++; return c; });

    Time("Unsafe (c++)", () => { long c = 0; Parallel.For(0, N, i => { c++; }); return c; });

    Time("lock", () =>
    {
        long c = 0; object gate = new object();
        Parallel.For(0, N, i => { lock (gate) { c++; } });
        return c;
    });

    Time("Interlocked", () =>
    {
        long c = 0;
        Parallel.For(0, N, i => { Interlocked.Increment(ref c); });
        return c;
    });

    Time("Thread-local + merge", () =>
    {
        long total = 0;
        Parallel.For(0, N, () => 0L, (i, state, local) => local + 1,
            local => Interlocked.Add(ref total, local));
        return total;
    });
}

Two details make this harness fairer than a naive one. The warm-up call keeps JIT compilation and thread-pool start-up out of the first measurement, and running three rounds exposes how much a single timing can vary.

2. What counter++ Really Does

The statement looks like one operation. It is three: read the current value, add one, write the result back.

If Thread A and Thread B both read 5 before either writes, both write 6. Two increments ran and the counter moved by one.

There is a second detail worth knowing. In the lambda i => { counter++; }, the variable counter is captured, so the compiler moves it into a field on a compiler-generated closure object. Every iteration, on every thread, reads and writes the same field on that same object. The sharing is invisible in the source code, which is part of why this bug is so easy to write.

3. The Evidence: Thirteen Wrong Answers

I collected 13 runs of the unsafe version across three versions of the test program. The expected value was 100,000. The results, sorted:

1,020

3,895

4,695

5,311

5,676

6,425

6,531

7,433

17,825

22,577

25,027

30,306

32,690

None came close to 100,000, and none repeated. Between roughly 67% and 99% of the increments were lost, with a median around 6,500. Nothing threw an exception and nothing was logged. A wrong answer that changes on every run is the signature of a race condition, and it is also why a test that passes once proves nothing.

4. The Timings

For the Release build, the first round of each version includes compilation costs, so this comparison uses rounds 2 and 3.

Version

Result

Time (ms)

Sequential loop

correct

0.033 to 0.034

Thread-local + merge

correct

0.277 to 0.287

Unsafe counter++

wrong

0.751 to 1.093

Interlocked

correct

2.212 to 2.667

lock

correct

9.023 to 22.809

This is one machine, one process, and a run started from Visual Studio, so the ordering and rough size of the gaps are meaningful, while the decimals are not.

5. The Surprise: The Plain Loop Won

The sequential loop beat every parallel version. The best correct parallel approach, thread-local accumulation, was roughly 8 to 9 times slower. Interlocked was around 65 to 80 times slower, and lock roughly 270 to 670 times slower.

The sequential loop finishing 100,000 iterations in about 33 microseconds is so fast that the compiler very likely optimized it heavily, so read that number as "this work costs almost nothing". That is exactly the point. Going parallel adds fixed costs: partitioning the range, scheduling tasks, coordinating threads, and merging results. When each unit of work is tiny, those costs dominate.

Parallelism pays off only when each item of work is heavy enough to outweigh the overhead, such as a database call, network I/O, or a large computation. Incrementing an integer is nowhere near that threshold.

A second oddity: the unsafe version, which gives wrong answers, was about 3 to 4 times slower than the correct thread-local version. A plausible explanation is that every core keeps competing for the same memory location, but I did not measure the cause, so treat it as a hypothesis. The measured fact is that sharing one variable was slower even before correctness is considered.

6. Four Correct Approaches, Under the Hood

6.1 lock

csharp

object gate = new object();
Parallel.For(0, N, i => { lock (gate) { counter++; } });

For an ordinary object, lock compiles to Monitor.Enter and Monitor.Exit wrapped in a try/finally. Only one thread at a time can be inside the block, so the read, add, and write can no longer interleave. Other threads that arrive while it is held must wait. Acquiring and releasing the lock also ensures that other threads see the updated value, which plain unsynchronized reads and writes do not guarantee.

Two practices matter. Lock on a private, readonly object, never on this or a public object, because outside code could lock the same one and create deadlocks. And keep the protected section as small as possible, since everything inside it is serialized.

lock was also the least consistent in my runs, from about 9 ms to 51 ms across rounds, which fits a mechanism that depends on how often threads collide.

6.2 Interlocked

csharp

Parallel.For(0, N, i => { Interlocked.Increment(ref counter); });

Interlocked performs the read-add-write as one indivisible operation, and on common hardware this typically maps to a single atomic processor instruction. Threads do not wait on a lock, and the operation also makes its result visible to other threads. In my Release runs it was about 3 to 10 times faster than lock.

Its limit matters more than its speed: it protects one variable and one operation. For a custom update, such as keeping the largest value seen so far, the standard pattern is a compare-and-swap loop:

csharp

static void AtomicMax(ref int target, int value)
{
    int current;
    do
    {
        current = Volatile.Read(ref target);
        if (value <= current) return;           // nothing to do
    }
    while (Interlocked.CompareExchange(ref target, value, current) != current);
}

CompareExchange writes the new value only if the variable still holds the value you read. If another thread changed it in the meantime, the loop retries with fresh data. If an update involves several variables that must stay consistent with each other, Interlocked cannot help, and lock is the right tool.

6.3 Thread-local state, merged once

csharp

long total = 0;
Parallel.For(0, N,
    () => 0L,                                    // each task starts with its own counter
    (i, state, local) => local + 1,              // count privately
    local => Interlocked.Add(ref total, local)); // merge once per task

Each task counts into its own private variable and only combines results at the end. During the counting phase nothing is shared, so there is nothing to protect. This was the fastest correct parallel version in my runs.

One caution I did not test: if you build the same idea with an array of per-thread counters, adjacent elements can sit on the same cache line, and the cores can end up contending anyway, a problem known as false sharing. Keeping the local state in a genuinely private variable, as above, avoids it.

6.4 A PLINQ aggregate

csharp

long total = ParallelEnumerable.Range(0, N).Sum(i => 1L);

PLINQ's built-in aggregates manage partial results per partition and combine them for you, which is the thread-local idea with the bookkeeping done by the framework. I did not benchmark this variant, so I make no performance claim about it, but for sums, counts, and similar aggregates it is worth trying before hand-writing synchronization.

7. Misconceptions That Cause Real Bugs

"volatile makes it thread-safe." It does not make ++ atomic. volatile affects visibility and ordering of reads and writes, but an increment is still a separate read, add, and write, so updates can still be lost. It can only be applied to fields, so testing it needs a field:

csharp

static class Shared { public static volatile int Counter; }

Parallel.For(0, N, i => { Shared.Counter++; });   // still not atomic

You can confirm this by adding the variant to the harness above. I have not benchmarked it.

"A concurrent collection makes my logic thread-safe." Each individual operation on a ConcurrentDictionary is safe, but a check followed by an action is two operations:

csharp

// Racy: two threads can both pass the check, and the increment is also read-modify-write
if (!dict.ContainsKey(key)) dict[key] = 0;
dict[key]++;

// Safer: one atomic call
dict.AddOrUpdate(key, 1, (k, old) => old + 1);

Note that the update delegate in AddOrUpdate can be called more than once if threads collide, so it must not have side effects.

"No exception means no bug." Every unsafe run above finished without a single error.

8. When the Protected Code Needs await

lock cannot contain an await; the compiler rejects it. For asynchronous code, the equivalent tool is SemaphoreSlim:

csharp

private static readonly SemaphoreSlim _gate = new SemaphoreSlim(1, 1);

async Task UpdateAsync()
{
    await _gate.WaitAsync();
    try
    {
        await SaveAsync();      // awaiting inside the guarded section is allowed here
    }
    finally
    {
        _gate.Release();        // always release, even if SaveAsync throws
    }
}

Releasing inside finally is essential. Without it, one exception leaves the gate closed and every later caller waits forever.

9. Choosing a Strategy

Situation

Reach for

Work per item is tiny

Stay sequential and measure first

Totals or counts across a large loop

Thread-local state, or a PLINQ aggregate

One shared number, simple operation

Interlocked

Custom update of one variable

Interlocked.CompareExchange loop

Several values that must change together

lock

Protected code needs await

SemaphoreSlim

10. Measuring Without Fooling Yourself

The practices that made this experiment worth trusting:

For precise numbers, a purpose-built tool handles warm-up, repetition, and statistics for you. A BenchmarkDotNet version of this experiment, which you would run with dotnet run -c Release, looks like this:

csharp

using BenchmarkDotNet.Attributes;
using BenchmarkDotNet.Running;

public class CounterBenchmarks
{
    private const int N = 100_000;

    [Benchmark(Baseline = true)]
    public long Sequential() { long c = 0; for (int i = 0; i < N; i++) c++; return c; }

    [Benchmark]
    public long LockCounter()
    {
        long c = 0; object gate = new();
        Parallel.For(0, N, i => { lock (gate) { c++; } });
        return c;
    }

    [Benchmark]
    public long InterlockedCounter()
    {
        long c = 0;
        Parallel.For(0, N, i => { Interlocked.Increment(ref c); });
        return c;
    }

    [Benchmark]
    public long ThreadLocalCounter()
    {
        long total = 0;
        Parallel.For(0, N, () => 0L, (i, s, local) => local + 1,
            local => Interlocked.Add(ref total, local));
        return total;
    }
}

// In Program.cs:
// BenchmarkRunner.Run<CounterBenchmarks>();

11. Applying This to Batch Processing

A payroll-style batch looks like a good candidate for parallelism, since each employee's work involves database operations that are far heavier than incrementing a number. The same discipline applies. Each thread needs its own connection and its own transaction, never shared, because most database connection objects are not safe for concurrent use. The real speed-up depends on how heavy each record's work is and how the database handles concurrent writes, so it will not be a clean division by the number of threads. Measure the sequential version first, then the parallel one, and keep whichever is both faster and correct.

Takeaway

An unsynchronized shared counter lost between roughly 67% and 99% of its updates, differently on every run, without a single error. Interlocked is the right tool for one shared number, lock is for updates spanning several values, SemaphoreSlim replaces lock when await is involved, and the fastest correct option was to avoid sharing at all. The most useful result was the least expected one: for tiny work, every parallel version lost to a plain loop. Before parallelizing anything, measure the sequential version, then let the evidence decide.