When a production incident occurs, an engineering team needs more than an investigation result. The result must reach the people responsible for resolving the issue, become part of the incident record, and remain available for follow-up work.

AWS DevOps Agent can investigate operational issues and emit lifecycle events through Amazon EventBridge. By connecting EventBridge to AWS Lambda and Jira, teams can automatically create tickets when investigations complete, fail, or require attention.

This integration provides a practical way to connect automated incident investigation with an existing engineering workflow. EventBridge handles event filtering and routing, Lambda transforms the event into a Jira issue, and Jira becomes the place where engineers track ownership, remediation, and resolution.

The implementation requires careful handling of duplicate events, authentication, retries, and investigation details. A working integration should do more than create tickets: it should produce useful incident records without flooding Jira with duplicate or incomplete issues.

Understand the Integration Architecture

The integration consists of four main components.

The event flow looks like this:

  1. AWS DevOps Agent creates or updates an investigation.

  2. The agent emits an event to the default EventBridge event bus.

  3. An EventBridge rule matches the event type and routes it to Lambda.

  4. Lambda extracts the investigation metadata and prepares a Jira issue.

  5. Lambda submits the issue to Jira using an authenticated HTTPS request.

  6. The function records the outcome so failures can be diagnosed and duplicate processing can be controlled.

AWS DevOps Agent events use the source aws.aidevops. The event's detail-type identifies the lifecycle transition, such as Investigation Completed, Investigation Failed, or Mitigation Failed.

One important detail is that an event can contain identifiers for an investigation summary without containing the actual summary text. If the Jira ticket needs the complete findings, the integration must retrieve those records separately.

Step 1: Choose Which Events Should Create Jira Issues

Do not create a new Jira ticket for every lifecycle transition by default. An investigation can emit several events as it moves from creation to completion, and a mitigation can produce additional events.

Start with a narrow event pattern. For example, the following EventBridge rule matches completed and failed investigations and mitigations.

{
  "source": ["aws.aidevops"],
  "detail-type": [
    "Investigation Completed",
    "Investigation Failed",
    "Mitigation Failed"
  ]
}

This pattern focuses on outcomes that may require engineering follow-up.

You can create the rule in the EventBridge console by selecting the default event bus, choosing an event pattern, and adding a Lambda target. You can also manage the rule through infrastructure as code.

A useful production policy is to define event-to-ticket behavior explicitly:

Event

Suggested Jira behavior

Investigation Completed

Create a ticket if findings require follow-up

Investigation Failed

Create an investigation-failure ticket

Investigation Timed Out

Create a ticket or alert for triage

Mitigation Failed

Create a remediation-failure ticket

Mitigation Completed

Update an existing issue or record the outcome

The final column is a design recommendation, not an automatic behavior supplied by AWS. Your team must implement the desired mapping.

For a simple first version, creating tickets only for failed investigations and mitigations can reduce noise. After validating the workflow, expand it to successful investigations that identify actionable findings.

Step 2: Prepare Jira Authentication

For Jira Cloud, the REST API can create issues through the issue-creation endpoint. The integration needs a Jira account with permission to create issues in the target project.

Use a dedicated integration account rather than an individual engineer's credentials. Generate an appropriate API token and store it in AWS Secrets Manager.

A secret can contain the following fields:

{
  "baseUrl": "https://your-company.atlassian.net",
  "email": "[email protected]",
  "apiToken": "REPLACE_WITH_API_TOKEN",
  "projectKey": "OPS"
}

Replace the example values with your Jira Cloud configuration. The token shown here is a placeholder; do not commit a real token to source control or embed it in Lambda environment variables as plaintext.

The Lambda execution role needs permission to retrieve the specific secret. Restrict access to the required secret ARN instead of granting unrestricted access to all secrets.

Also configure network access. A Lambda function in a private VPC may need an appropriate outbound route to reach Jira Cloud. Without connectivity, the function can fail even when its credentials and event pattern are correct.

Step 3: Create the Lambda Function

The following Python example accepts an EventBridge event, validates its basic structure, retrieves Jira credentials from Secrets Manager, and creates a Jira issue.

It uses the Jira Cloud REST API v3, including Atlassian Document Format for the description. The example creates one ticket per invocation; it does not yet provide cross-invocation deduplication.

Configure these Lambda environment variables:

Use a supported Python runtime and grant the function permission to read the configured secret.

import base64
import json
import os
import urllib.error
import urllib.request

import boto3

secrets = boto3.client("secretsmanager")

SECRET_ARN = os.environ["JIRA_SECRET_ARN"]
ISSUE_TYPE = os.environ.get("JIRA_ISSUE_TYPE", "Task")


def load_jira_config():
    response = secrets.get_secret_value(
        SecretId=SECRET_ARN
    )
    return json.loads(response["SecretString"])


def paragraph(text):
    return {
        "type": "paragraph",
        "content": [
            {
                "type": "text",
                "text": str(text)
            }
        ]
    }


def create_jira_issue(config, event):
    detail = event.get("detail", {})
    metadata = detail.get("metadata", {})
    data = detail.get("data", {})

    event_type = event.get(
        "detail-type", "Unknown event"
    )

    task_id = metadata.get("task_id", "unknown")
    execution_id = metadata.get(
        "execution_id", "not assigned"
    )
    agent_space_id = metadata.get(
        "agent_space_id", "unknown"
    )

    priority = data.get("priority", "MEDIUM")
    status = data.get("status", "UNKNOWN")
    summary_record_id = data.get(
        "summary_record_id", "not available"
    )

    summary = (
        f"AWS DevOps Agent: {event_type} "
        f"(task {task_id})"
    )

    description = {
        "type": "doc",
        "version": 1,
        "content": [
            paragraph(
                "An AWS DevOps Agent lifecycle event "
                "was received."
            ),
            paragraph(f"Event: {event_type}"),
            paragraph(f"Task ID: {task_id}"),
            paragraph(f"Execution ID: {execution_id}"),
            paragraph(f"Agent Space ID: {agent_space_id}"),
            paragraph(f"Priority: {priority}"),
            paragraph(f"Status: {status}"),
            paragraph(
                f"Summary record ID: {summary_record_id}"
            ),
            paragraph(
                f"AWS account: {event.get('account', 'unknown')}"
            ),
            paragraph(
                f"AWS Region: {event.get('region', 'unknown')}"
            ),
            paragraph(
                f"Event ID: {event.get('id', 'unknown')}"
            )
        ]
    }

    payload = {
        "fields": {
            "project": {
                "key": config["projectKey"]
            },
            "issuetype": {
                "name": ISSUE_TYPE
            },
            "summary": summary[:255],
            "description": description,
            "labels": [
                "aws-devops-agent",
                "automated-investigation"
            ]
        }
    }

    credentials = (
        f"{config['email']}:{config['apiToken']}"
    ).encode("utf-8")

    auth = base64.b64encode(
        credentials
    ).decode("ascii")

    url = (
        config["baseUrl"].rstrip("/")
        + "/rest/api/3/issue"
    )

    request = urllib.request.Request(
        url,
        data=json.dumps(payload).encode("utf-8"),
        headers={
            "Authorization": f"Basic {auth}",
            "Accept": "application/json",
            "Content-Type": "application/json"
        },
        method="POST"
    )

    with urllib.request.urlopen(
        request, timeout=10
    ) as response:
        return json.loads(
            response.read().decode("utf-8")
        )


def lambda_handler(event, context):
    event_type = event.get("detail-type", "")

    allowed_events = {
        "Investigation Completed",
        "Investigation Failed",
        "Mitigation Failed"
    }

    if event.get("source") != "aws.aidevops":
        raise ValueError("Unexpected event source")

    if event_type not in allowed_events:
        return {
            "statusCode": 200,
            "message": "Event ignored"
        }

    if not event.get("detail", {}).get("metadata", {}).get(
        "task_id"
    ):
        raise ValueError("Missing investigation task ID")

    config = load_jira_config()

    try:
        issue = create_jira_issue(config, event)
    except urllib.error.HTTPError as exc:
        error_body = exc.read().decode(
            "utf-8", errors="replace"
        )
        print(
            json.dumps({
                "eventId": event.get("id"),
                "httpStatus": exc.code,
                "error": error_body[:1000]
            })
        )
        raise

    return {
        "statusCode": 200,
        "issueKey": issue.get("key"),
        "issueId": issue.get("id"),
        "eventId": event.get("id")
    }

The code intentionally creates a descriptive ticket from event metadata rather than claiming that the full investigation findings are included. The summary_record_id is an identifier, not the summary itself.

The code also logs a limited portion of Jira's error response to help diagnose API failures. Review logging policies before increasing the amount of response data retained, especially if incident information could contain sensitive details.

Validate the Jira project fields

Jira projects can have different required fields and issue types. The example assumes the project accepts Task and allows the supplied fields.

If issue creation fails with a validation error, inspect the Jira response and project configuration. You may need to select another issue type or provide required fields such as a component, custom field, or additional project-specific value.

Do not resolve these failures by granting the integration account unnecessary administrative permissions. Give it only the access required to create and, if applicable, update the intended issues.

Step 4: Add the EventBridge Target and Permissions

After creating the Lambda function, attach it as the target of the EventBridge rule.

EventBridge needs permission to invoke the function. When configuring the target through the console, the console can help establish the required resource-based permission. For infrastructure-as-code deployments, explicitly configure the Lambda permission for the EventBridge rule.

The permission should identify the relevant EventBridge rule as the allowed invocation source. Avoid allowing arbitrary principals to invoke the function.

The Lambda execution role and the Lambda resource policy serve different purposes:

Keeping these permissions separate makes the integration easier to review and troubleshoot.

Step 5: Include the Actual Investigation Findings

A Jira ticket containing only an event type, task ID, and status provides traceability, but it may not provide enough context for an engineer to resolve the incident.

AWS DevOps Agent events can include data.summary_record_id. This identifies a summary record when one is available, but the event does not carry the summary text itself. The field can also be absent, including for a completed investigation that produced no summary.

To include the findings, extend the integration to retrieve the corresponding journal record through the AWS DevOps Agent API before creating the ticket.

The retrieval process needs the relevant agent space ID, execution ID, record type, and summary record ID. Investigation summaries use the investigation_summary_md record type, while mitigation summaries use mitigation_summary_md.

A robust implementation should:

  1. Confirm that the event contains the identifiers needed for retrieval.

  2. Retrieve the journal records using the supported AWS DevOps Agent API operation.

  3. Select the record whose ID matches summary_record_id.

  4. Add the retrieved findings to the Jira description.

  5. Handle missing records and temporary API failures without silently losing the event.

The Lambda execution role must have the permissions required to retrieve these records. Do not assume the event itself contains enough information to reconstruct the full investigation.

If the summary is unavailable, the integration can still create a ticket with the available metadata and mark the findings as unavailable. This preserves traceability without inventing incident details.

Step 6: Prevent Duplicate Jira Issues

Event-driven systems must account for retries and duplicate processing. EventBridge and Lambda can deliver or process an event more than once, so a successful Jira issue creation does not guarantee that the same event will never be processed again.

Without deduplication, a temporary timeout could cause Lambda to retry an invocation even if Jira already created the issue. The result can be multiple tickets for one investigation event.

Use the EventBridge event ID as part of a durable idempotency strategy. One approach is to store processed event IDs in DynamoDB before or alongside the Jira operation, with a conditional write that prevents two invocations from claiming the same event simultaneously.

A production workflow needs more than a simple “processed” flag. Consider these states:

There is an important edge case: Jira may create an issue successfully while Lambda times out before saving the resulting issue key. A retry can then create a duplicate unless the workflow can recover the earlier result.

For stronger protection, use a stable mapping between the event ID and Jira issue key, and consider an external identifier or custom Jira field that allows the integration to find an issue created during an earlier attempt. Make the lookup and creation process concurrency-safe.

For workflows that need to group several lifecycle events into a single incident ticket, use a stable investigation identifier such as the agent space and task ID as the grouping key. That requires an explicit update-or-create workflow rather than treating every event as an independent ticket.

Step 7: Configure Retries, Timeouts, and Monitoring

A Jira integration depends on two independently operated systems: AWS event processing and the Jira API. Temporary network failures, API throttling, expired credentials, and Jira service interruptions can affect delivery.

Configure Lambda timeouts based on the expected API response time, and set EventBridge target retry behavior according to the operational requirements. For events that must not be lost, configure an appropriate dead-letter queue and monitor it.

Classify failures before deciding whether to retry:

Add CloudWatch metrics and alarms for Lambda errors, throttling, duration, and dead-letter queue depth. Log the EventBridge event ID and resulting Jira issue key so operators can correlate an AWS event with its ticket.

Avoid logging API tokens, authorization headers, or unnecessary sensitive incident data.

Summary

Amazon EventBridge and AWS Lambda provide a flexible way to connect AWS DevOps Agent lifecycle events with Jira. EventBridge filters investigation and mitigation events, Lambda converts the event into a Jira issue, and the Jira record gives engineering teams a place to manage follow-up work.

A basic integration can create tickets using the task ID, event type, priority, status, and execution metadata. To include actual investigation findings, retrieve the corresponding journal record separately because the event contains a summary record identifier rather than the summary text.

For production use, prioritize idempotency, least-privilege permissions, retry handling, useful monitoring, and a clear policy for deciding when to create or update tickets. These controls turn a simple webhook-style connection into a reliable incident-management workflow.