Introduction

AWS DataSync is designed to move large amounts of data between AWS storage services, on-premises storage, and other supported locations. Once a DataSync task becomes part of a production workflow, monitoring becomes important for answering simple operational questions:

  • Is the transfer running?

  • How much data has moved?

  • Is the task failing?

  • Are files being skipped?

  • Is the transfer progressing normally?

  • Did the destination receive and verify the expected data?

A common approach is to build custom Lambda functions, poll the DataSync API, publish custom CloudWatch metrics, and then create alarms around those metrics.

That can work, but it is often unnecessary.

AWS DataSync already publishes service metrics to Amazon CloudWatch automatically. These metrics are available in the AWS/DataSync namespace and are emitted at five-minute intervals. DataSync also provides task execution details, CloudWatch Logs, and task reports for deeper investigation.

The better starting point is therefore to use the telemetry DataSync already provides and add custom metrics only when the built-in information cannot answer a specific business or application question.

What DataSync Already Provides

DataSync provides several monitoring mechanisms:

AWS DataSync
     |
     +---- CloudWatch Metrics
     |
     +---- CloudWatch Logs
     |
     +---- Task Execution Details
     |
     +---- Task Reports

Each serves a different purpose.

Monitoring method

Best for

CloudWatch Metrics

Transfer progress and performance

CloudWatch Logs

Errors and detailed transfer activity

DescribeTaskExecution

Current execution state and counters

Task Reports

File/object-level transfer details

CloudWatch Alarms

Alerting on measurable conditions

Dashboards

Operational visibility

AWS recommends monitoring DataSync transfers and supports CloudWatch metrics, task reports, and CloudWatch Logs for this purpose.

DataSync CloudWatch Metrics

DataSync publishes metrics under:

AWS/DataSync

The metrics include dimensions such as:

TaskId
AgentId

The available metrics depend partly on the DataSync task mode.

Examples include:

  • BytesTransferred

  • BytesCompressed

  • BytesWritten

  • FilesTransferred

  • FilesSkipped

  • FilesDeleted

  • FilesVerified

  • BytesPreparedSource

  • BytesPreparedDestination

  • FilesPreparedSource

  • FilesPreparedDestination

AWS currently publishes these metrics automatically to CloudWatch at five-minute intervals, and the metrics are retained for 15 months.

This is important because you do not need to write code that repeatedly calls the DataSync API simply to produce basic transfer metrics.

Start With the Built-In Metrics

Suppose you have a nightly DataSync task.

A basic operational dashboard could show:

DataSync Task
------------------------------
Status
Bytes Transferred
Files Transferred
Files Skipped
Files Verified
Transfer Throughput

The architecture can remain simple:

DataSync
   |
   v
CloudWatch Metrics
   |
   +---- Dashboard
   |
   +---- Alarm
   |
   +---- Operations Team

There is no Lambda function in this design.

There is also no custom metric publishing pipeline.

Why Custom Metrics Are Often Unnecessary

A custom monitoring implementation might look like this:

EventBridge
    |
    v
Lambda
    |
    v
DescribeTaskExecution
    |
    v
PutMetricData
    |
    v
CloudWatch

That introduces additional components.

You now need to manage:

  • Lambda code

  • IAM permissions

  • Error handling

  • Retries

  • Scheduling

  • API throttling

  • Metric naming

  • Metric dimensions

  • Lambda monitoring

  • Deployment

  • Maintenance

If DataSync already exposes the required measurement, this additional infrastructure does not provide much value.

A simpler design is:

DataSync
    |
    v
CloudWatch
    |
    +---- Alarm
    +---- Dashboard

Use custom metrics when you actually need information that DataSync does not provide.

Monitor Transfer Progress

BytesTransferred and file-related metrics can help determine whether a transfer is progressing.

For example:

10:00  100 GB
10:05  145 GB
10:10  190 GB
10:15  240 GB

The increasing values indicate transfer activity.

However, avoid treating every DataSync metric as an exact accounting value.

AWS notes that FilesTransferred can differ from estimated file counts in some circumstances and that its behavior is implementation-specific for some location types. It should not automatically be treated as an exact indication of everything transferred.

This is an important distinction between operational telemetry and authoritative reconciliation data.

Detect Failed Transfers

A production monitoring system should detect failures rather than simply display throughput.

DataSync execution information includes result and status information that can be retrieved using DescribeTaskExecution. When an execution fails, the result can include error details that help identify the problem.

For example:

Task
 |
 +---- RUNNING
 |
 +---- SUCCESS
 |
 +---- ERROR
 |
 +---- CANCELED

The exact states and response fields should be handled according to the current DataSync API.

For alerting, the objective is straightforward:

Transfer failure
      |
      v
CloudWatch / Event-driven alert
      |
      v
Operations team

Use CloudWatch Logs for Errors

Metrics are useful for determining that something is wrong.

Logs help explain what went wrong.

AWS recommends configuring DataSync tasks to log at least basic information such as transfer errors. Enhanced mode can automatically send task logs to a /aws/datasync log group when the required service-linked role and configuration are available.

A typical architecture is:

DataSync Task
     |
     v
CloudWatch Logs
     |
     +---- Error investigation
     +---- Log Insights
     +---- Operational troubleshooting

This is generally more useful than creating a custom metric for every possible error message.

Use Task Reports for File-Level Investigation

CloudWatch metrics tell you about the transfer as a whole.

Sometimes you need to know exactly which objects failed.

This is where DataSync task reports are useful.

Task reports can provide information about files, objects, and directories that DataSync attempted to transfer, skip, verify, or delete. Reports can be configured as summary-only or with more detailed information.

For example:

Task Execution
      |
      v
Task Report
      |
      +---- file-a.csv  SUCCESS
      +---- file-b.csv  SUCCESS
      +---- file-c.csv  ERROR
      +---- file-d.csv  SKIPPED

This provides a much better troubleshooting path than creating a custom metric such as:

FailedFiles = 1

A metric tells you that a problem exists.

A task report can help identify the affected object.

Monitor Different Layers Separately

A useful monitoring design separates three layers.

Layer 1: Execution

Questions:

  • Did the task start?

  • Is it running?

  • Did it finish?

  • Did it fail?

Layer 2: Transfer

Questions:

  • How many bytes moved?

  • How many files moved?

  • How many files were skipped?

  • How much data was written?

Layer 3: Data Validation

Questions:

  • Was the data verified?

  • Which files failed verification?

  • Which objects were skipped?

  • What was deleted?

The resulting architecture looks like:

                DataSync
                   |
       +-----------+-----------+
       |           |           |
       v           v           v
   Execution    Transfer    Validation
       |           |           |
       v           v           v
   Status       Metrics    Task Reports
       |           |           |
       +-----------+-----------+
                   |
                   v
              Operations

Create a Practical CloudWatch Dashboard

A dashboard does not need dozens of graphs.

For most DataSync workloads, start with a small set of signals.

Transfer Volume

Use:

BytesTransferred
BytesWritten

Transfer Activity

Use:

FilesTransferred
FilesSkipped

Verification

Use:

FilesVerified

Operational Status

Use execution information and logs to investigate failed or abnormal executions.

The goal is to make the dashboard answer:

Is this transfer healthy right now?

rather than displaying every metric available.

Example CloudWatch Metric Query

The AWS CLI can retrieve DataSync metrics from CloudWatch.

A simplified example is:

aws cloudwatch get-metric-statistics \
  --namespace AWS/DataSync \
  --metric-name BytesTransferred \
  --dimensions Name=TaskId,Value=task-0123456789abcdef0 \
  --statistics Sum \
  --period 300 \
  --start-time 2026-09-28T00:00:00Z \
  --end-time 2026-09-28T02:00:00Z

This approach uses the metric DataSync already publishes.

You do not need to call DescribeTaskExecution and republish the result as a custom CloudWatch metric simply to graph transfer volume.

Use DescribeTaskExecution for Detailed State

CloudWatch metrics are not the only source of operational information.

The DataSync API provides DescribeTaskExecution, which returns information about a particular task execution. This can include transfer counters, estimated transfer information, status, and result details.

For example:

aws datasync describe-task-execution \
  --task-execution-arn \
  arn:aws:datasync:region:account-id:task/task-id/execution/execution-id

This is particularly useful during troubleshooting.

A practical workflow is:

CloudWatch Alarm
      |
      v
Identify failed task
      |
      v
DescribeTaskExecution
      |
      v
CloudWatch Logs
      |
      v
Task Report

Each step adds more detail.

When Custom Metrics Actually Make Sense

There are legitimate reasons to create custom metrics.

For example, suppose the business requirement is:

"Alert when less than 98% of expected objects arrive."

DataSync may provide the transfer counters, but the expected object count could come from your application or inventory system.

You could then calculate:

TransferCompletionRate =
TransferredObjects / ExpectedObjects * 100

That is a workload-specific metric.

Another example:

BusinessDataFreshness

This could represent:

CurrentTime - LastSuccessfulDataSync

That is different from a raw DataSync metric because it represents a business-level condition.

The principle is:

Use native metrics for service behavior and custom metrics for workload-specific meaning.

Common Mistakes

Building Lambda Polling Before Checking Native Metrics

This adds infrastructure unnecessarily.

Treating Every Metric as an Exact Data Reconciliation Source

Operational metrics can have service-specific semantics.

Using Metrics for File-Level Troubleshooting

Use task reports and logs when you need to identify individual objects.

Creating Too Many Alarms

An alarm for every metric can create noise.

Focus on conditions that require action.

Ignoring Task Mode Differences

Not every DataSync metric is available in every task mode.

Monitoring Only Transfer Volume

A transfer can move data while still producing skipped, verification, or destination-side problems.

Best Practices

  1. Start with the native AWS/DataSync CloudWatch metrics.

  2. Use dashboards for high-level operational visibility.

  3. Use CloudWatch Logs for error investigation.

  4. Use DescribeTaskExecution when detailed execution state is required.

  5. Use task reports for file- and object-level analysis.

  6. Create alarms around actionable failure or freshness conditions.

  7. Avoid publishing a custom metric when an equivalent native metric already exists.

  8. Treat custom metrics as business or application-level measurements.

  9. Review task mode before selecting metrics.

  10. Keep monitoring infrastructure simpler than the workload it monitors.

Advantages and Disadvantages

Advantages

Disadvantages

Less monitoring code

Native metrics may not cover every business requirement

No polling Lambda required for basic metrics

Metrics are published at five-minute intervals

Lower operational complexity

Some metrics have task-mode limitations

Native CloudWatch integration

File-level details require logs or reports

Easy dashboard integration

Custom business calculations may still require additional logic

A Simple Production Architecture

For many DataSync workloads, the following design is sufficient:

                    AWS DataSync
                         |
          +--------------+--------------+
          |              |              |
          v              v              v
     CloudWatch       CloudWatch      Task Reports
      Metrics           Logs             |
          |              |               |
          v              v               v
      Dashboard       Errors        File Analysis
          |
          v
       Alarms
          |
          v
     Operations Team

There is no custom Lambda metric pipeline unless the workload actually requires one.

Production Checklist

[ ] DataSync metrics are visible in CloudWatch
[ ] Correct TaskId dimensions are being monitored
[ ] Transfer volume is monitored
[ ] Transfer failures are detectable
[ ] CloudWatch Logs are enabled where required
[ ] Task reports are configured for workloads that need file-level detail
[ ] Dashboard focuses on actionable signals
[ ] Alarms have clear operational owners
[ ] Task mode differences have been reviewed
[ ] Custom metrics are used only for workload-specific requirements
[ ] Monitoring has been tested with a failed transfer

Summary

AWS DataSync already provides a substantial monitoring foundation through CloudWatch metrics, CloudWatch Logs, task execution details, and task reports.

The DataSync metrics in the AWS/DataSync namespace are automatically published to CloudWatch, which means many production workloads do not need a Lambda function that repeatedly polls DataSync and publishes custom metrics.

Use the tools according to the question you need to answer:

"Is the transfer progressing?"
        |
        v
CloudWatch Metrics

"Why did it fail?"
        |
        v
CloudWatch Logs / Execution Details

"Which files failed?"
        |
        v
Task Reports

"Does this meet my business requirement?"
        |
        v
Custom Metric, if necessary

This approach keeps the monitoring architecture smaller while still providing the information needed for day-to-day operations and troubleshooting.

The key principle is simple: use AWS DataSync's native telemetry first, and build custom monitoring only where the native signals stop being sufficient.