Skip to content

Terraform setup - #37

Merged
KaseyW31 merged 6 commits into
qafrom
add-terraform
Sep 23, 2026
Merged

KaseyW31 merged 6 commits into
qafrom
add-terraform

Conversation

@KaseyW31

@KaseyW31 KaseyW31 commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

Sets up Terraform and an alarm. Also makes some changes to the Dockerfile so that tests can run (issues with deprecations). I don't know much about Ruby/this setup so other changes may be needed

Update: the alarm was deleted somehow so I removed the import and updated the terraform plan output

Questions for reviewers:

  1. An alarm currently exists with the same name in Cloudwatch. Would it be preferred to delete the original and use this config to define a new one, or modify the existing one and import it?
  2. The existing alarm is set on a log filter on the metric SCSBusterErrorAlarm. In this PR, an new metric SCSBusterError and filter (identical except for name) are defined to keep with the current naming convention. Would using the original metric be preferred?
  3. Assuming that we currently don't want an alarm for QA - could either keep the current modularity to allow for a QA setup if wanted, or I could condense everything into one module

terraform plan output:

Terraform used the selected providers to generate the following execution plan. Resource actions are indicated with the following symbols:
  + create

Terraform will perform the following actions:

  # aws_cloudwatch_log_metric_filter.log_error will be created
  + resource "aws_cloudwatch_log_metric_filter" "log_error" {
      + apply_on_transformed_logs = (known after apply)
      + id                        = (known after apply)
      + log_group_name            = "/ecs/scsbuster-production-tf"
      + name                      = "SCSBusterLogError"
      + pattern                   = "{ $.level = \"FATAL\" }"
      + region                    = "us-east-1"

      + metric_transformation {
          + name      = "SCSBusterLogError"
          + namespace = "LogMetrics"
          + unit      = "None"
          + value     = "1"
        }
    }

  # aws_cloudwatch_metric_alarm.log_error_alarm will be created
  + resource "aws_cloudwatch_metric_alarm" "log_error_alarm" {
      + actions_enabled                       = true
      + alarm_actions                         = [
          + "arn:aws:sns:us-east-1:946183545209:research-catalog-team-alarms-production",
        ]
      + alarm_description                     = "Triggered when there's 1 or more fatal error log to SCSBuster within 5 minutes."
      + alarm_name                            = "SCSBusterLogErrorAlarm"
      + arn                                   = (known after apply)
      + comparison_operator                   = "GreaterThanOrEqualToThreshold"
      + datapoints_to_alarm                   = 1
      + evaluate_low_sample_count_percentiles = (known after apply)
      + evaluation_periods                    = 1
      + id                                    = (known after apply)
      + metric_name                           = "SCSBusterLogError"
      + namespace                             = "LogMetrics"
      + period                                = 300
      + region                                = "us-east-1"
      + statistic                             = "Sum"
      + tags_all                              = (known after apply)
      + threshold                             = 1
      + treat_missing_data                    = "notBreaching"
    }

Plan: 2 to add, 0 to change, 0 to destroy.

@7emansell 7emansell left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  1. Yes, we want to import the existing alarm and then modify it
  2. No, let's go ahead with the correct naming convention (the metric should be named SomethingError and the alarm should be named SomethingErrorAlarm)
  3. Yes, let's just keep it on prod for now

Comment thread terraform/production/alarms.tf Outdated
import {
to = aws_cloudwatch_metric_alarm.scsbuster_error_alarm
id = "SCSBusterErrorAlarm"
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Neat. I was not previously aware of import blocks. This seems maybe better than the import command since it documents that this was previously managed outside of terraform and how we imported it (because I always have to look up the right terraform import command).

@KaseyW31

KaseyW31 commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor Author

I just saw that the original alarm (SCSBusterErrorAlarm) was deleted ~3 weeks ago but I have no recollection of doing so... I'll see if I can figure out how it happened but in the meantime, I removed the import block so that a new alarm (with updated name) will be created.

I don't see any fatal error logs in Cloudwatch in the time without alarm coverage, but ideally this PR can be merged in soon to reduce further downtime

@danamansana danamansana left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lock files should be gitignored and not included in commit, otherwise looks good

@KaseyW31
KaseyW31 merged commit e3b9637 into qa Sep 23, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants