HTTP 502 randomly occuring

:hugs: Please help fill in this template with all the details to help others help you more efficiently. Use formatting blocks for code, config, logs and ensure to remove sensitive data.

Problem to solve

We have a self hosted GitLab instance, not in a docker container, running on an internal-only Ubuntu 24.04LTS VM on Google Compute Engine.

We also have Uptime Kuma also hosted on a dedicated internal only VM on Google Compute Engine, monitoring the status of our various VMs across the organization, pinging them all once every 60 seconds. It pings the VM itself and GitLab separately. It checks the GitLab service is up and running by ensuring it receives HTTP status codes 200-299, as well as checking SSL certs, and other things. The VM itself is checked simply via ping to make sure it responds. This setup worked well for us for a number of years.

Over the last few months, we have noticed that GitLab - and only GitLab - will suddenly spit out HTTP status code 502 to our Kuma instance for 1-2 minutes (one or two “pings”). The VM itself remains solid with an expected uptime (currently at 17 days). The time this happens are unprompted and seemingly random. A screenshot below showing a brief history. I know

We initially thought it might have been a bad GitLab update, but several updates have been installed since then, and now we’re thinking it’s potentially something wrong with our infrastructure. Users have reported that just as randomly and infrequently, GitLab pages will partially load, requiring a refresh, so we’re wondering if this issue is related to that one.

However we’re trying to pinpoint what it could possibly be, so some guidance would be appreciated. It’s infrequent and random enough (sometimes months between occurrences) to make troubleshooting this difficult, so a guidance on what logs to review would hopefully point us in the right direction.

In terms of system specifications, I’m fairly confident we meet the recommended requirements. There’s <35 engineers on this instance and they start/finish work at any time between 08:00 - 20:00. lscpu displays the following, and we have 16GB RAM with an 8GB page file, and are currently on GitLab v19.2.0:

Architecture:                x86_64
  CPU op-mode(s):            32-bit, 64-bit
  Address sizes:             46 bits physical, 48 bits virtual
  Byte Order:                Little Endian
CPU(s):                      4
  On-line CPU(s) list:       0-3
Vendor ID:                   GenuineIntel
  BIOS Vendor ID:            Google
  Model name:                Intel(R) Xeon(R) CPU @ 2.20GHz
    BIOS Model name:           CPU @ 2.0GHz
    BIOS CPU family:         1
    CPU family:              6
    Model:                   79
    Thread(s) per core:      2
    Core(s) per socket:      2
    Socket(s):               1
    Stepping:                0
    BogoMIPS:                4399.99
    Flags:                   fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc rep_good nopl xtopology nonstop_tsc
                              cpuid tsc_known_freq pni pclmulqdq ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdrand hypervisor lahf_lm abm 3dnowprefetch pti ssbd ibrs ibpb stibp fsg
                             sbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm rdseed adx smap xsaveopt arat md_clear arch_capabilities

If you need further information on our setup, please let me know.

Any changes made to /etc/gitlab/gitlab.rb that might have affects on performance? One thing that springs to mind is if the puma configuration has been changed from it’s defaults.

Nothing that I can think of? Some extra things that I’ve configured which might affect it;

/var/opt/gitlab has been symlinked to a non-boot disk drive.

gitlab_rails['backup_path'] has also been set to a mount point that’s not the boot disk drive.

gitlab_rails['backup_keep_time'] = 2592000

nginx has been configured, along with ssl certs

nginx['enable'] = true
nginx['client_max_body_size'] = '1000m'
nginx['redirect_http_to_https'] = true
nginx['redirect_http_to_https_port'] = 80

Other than standard/expected configurations like email and OAUTH settings, the file is largely commented out.

Sounds like load / resource pressure when the server is not responding.

  1. Correlate the timestamp from the 502 with production logs, and identify the area/service that is causing it.
  2. Add performance monitoring with metrics like CPU, memory, disk I/O, etc. and correlate against the 502 errors.
  3. Investigate on GCP side if an update from 4 to 8 CPU cores changes the behavior and errors are gone. If yes, there are resource consumers / full queue locks worth investigating.

I haven’t used Uptime Kuma, but would also look into its configuration to extract the HTTP error message next to the status code only. Can you share the Kuma configuration to better understand how these monitoring calls are performed?

One thing to add, I don’t use GCP, but I’ve had issues for example in IBM Cloud if the CPU type chosen for the instance isn’t good enough. I had an instance from their cheaper CPU pricing and would have issues sometimes when the instance was inaccessible. The one I had chosen was basically meant for low load, shared CPU almost whereby you are not guaranteed to have access to the resources all of the time.

Since it was a test instance I wasn’t that bothered about it, but something to look at, depending on what type of instance you’ve chosen within GCP.

Hi, thanks to both for answering.

Correlate the timestamp from the 502 with production logs, and identify the area/service that is causing it.

I had a look at /var/log/gitlab/gitlab-rails/production.log but there’s no timestamps on any of the lines, so it makes figuring this out very difficult. I do see a large number of warnings repeated as below throughout the file, though I don’t think it’s related.

GraphQL-Ruby encountered mismatched types in this query: `Boolean!` (at 5:7) vs. `Boolean` (at 48:3).
This will return an error in future GraphQL-Ruby versions, as per the GraphQL specification
Learn about migrating here: https://graphql-ruby.org/api-doc/2.6.3/GraphQL/Schema.html#allow_legacy_invalid_return_type_conflicts-class_method

GraphQL-Ruby encountered mismatched types in this query: `Boolean!` (at 17:7) vs. `Boolean` (at 48:3).
This will return an error in future GraphQL-Ruby versions, as per the GraphQL specification
Learn about migrating here: https://graphql-ruby.org/api-doc/2.6.3/GraphQL/Schema.html#allow_legacy_invalid_return_type_conflicts-class_method
  • Is there a way to enable time stamping on this file so that I can better check this in the future?
  • Is the above error definitely not something to worry about/fixable?

Add performance monitoring with metrics like CPU, memory, disk I/O, etc. and correlate against the 502 errors.

I’ve noticed that whenever there’s downtime, there’s a large CPU/Activity spike. This VM is a dedicated VM for GitLab.

Apologies, I thought production_json.log was in that URL anchor section. Use /var/log/gitlab/gitlab-rails/production_json.log for structured log formatting including time as attribute.

Use this finding to dig deeper into why CPU activity spikes (or, CPU gets saturated). It could be a resource constraint - many services waiting for the same lock, the Sidekiq queues are full and workers are busy (too many requests, background jobs, etc.).

Additional logs: gitlab-workhorse/current and nginx/gitlab_error.log for the 502 origin, and Puma logs for worker timeouts/OOM kills. production_json.log sorted by duration_s around the spike window will show what was hot.

No, this is an internal Ruby warning that needs to be fixed in Fix GraphQL-Ruby type mismatch deprecation warnings (#586994) · Issues · GitLab.org / GitLab · GitLab

Just to confirm - GitLab Runner is installed on separate VMs?

I also asked Claude about this topic and additional troubleshooting. Sharing here, untested.


With 4 vCPUs and 16 GB RAM, a 502 next to a CPU spike usually means Puma had no worker free to answer Kuma’s request. A 1-2 minute outage matches that: one or two 60 second checks, then recovery. Every layer keeps its own log, so it is worth working through them in this order.

1. Find out which layer returns the 502. The request chain is NGINX to Workhorse to Puma.

grep -i "502\|upstream" /var/log/gitlab/nginx/gitlab_error.log
tail -f /var/log/gitlab/gitlab-workhorse/current

connection refused means Puma was down or restarting. A timeout means Puma was up but stuck. These are different problems, so this step decides where to look next.

2. Check whether Puma is being restarted.

gitlab-ctl tail puma
dmesg -T | grep -i -E "oom|killed process"

With 16 GB an OOM kill is unlikely, but GitLab also restarts workers that cross a memory threshold through the memory watchdog, and a restart cycle looks exactly like a short 502 window. Also worth checking whether the 8 GB page file is actually being used during a spike, since swapping alone can push request latency past the timeout.

3. Sidekiq and Puma competing for 4 cores. This is my main suspicion. Sidekiq defaults to a concurrency of 20, and Puma auto-detects its worker count from the available CPU and memory. Check what you actually run:

ps -eo pid,comm,args | grep -E "puma|sidekiq"

If a batch of background jobs saturates all 4 cores, Puma gets no CPU time and NGINX returns 502 until the queue drains. That fits both the spike shape and the fact that it clears itself after a minute or two. Lowering sidekiq['concurrency'] to 10 is a cheap experiment that does not need a VM resize.

4. Check whether the spikes line up with scheduled work. Sidekiq logs are JSON with durations, so long jobs are easy to spot:

jq -r 'select(.duration_s > 60) | "\(.time) \(.class) \(.duration_s)"' \
  /var/log/gitlab/sidekiq/current | sort -k3 -n | tail -30

Repository housekeeping, container registry cleanup and project exports are the usual candidates. Your backups are worth a separate look. With backup_keep_time at 30 days and the backup path on its own mount, a backup run is heavy on both CPU and disk. Do the 502 timestamps line up with your backup schedule?

5. Look at disk I/O, not only CPU. You symlinked /var/opt/gitlab to a non-boot disk, and on GCE the IOPS and throughput limits of a persistent disk scale with its type and size. A disk that is large enough for the data can still be too small for the throughput GitLab wants, and once it throttles, everything above it stalls. Time spent waiting on I/O also shows up in most graphs as activity, so a spike does not necessarily mean the CPU was doing work.

iostat -x 5

Watch %iowait and await during a spike, and compare against the disk throughput and IOPS charts in GCP monitoring. Which disk type and size did you use for that mount?

6. Catch it live if you can. top -H during a spike tells you whether the CPU goes to puma, git, postgres or sidekiq, which removes most of the guesswork in a single observation. The bundled Prometheus and Grafana under Admin > Monitoring already track Puma saturation, queue depth and per-process memory, so there is nothing to install first.

Hi there,

So, reviewing the logs you pointed me to (thank you!), I’m also starting to agree that there’s a constraint somewhere, and now I’m trying to figure out the where and why part.

Further information on our setup; we don’t have Gitlab runners on the Gitlab VM itself. A separate VM is in place with a single GitLab Runner in a container to receive the jobs. When a job gets fed through the pipeline, that runner receives it and spins up a fresh, disposable “runner VM” to execute the job (tags determine whether it’s a general purpose VM or a high powered VM for larger projects). have CI runners that get spun up via a ‘disposable’ VM. (Effectively this)

I don’t have logs for all of the downtime events, so I’ll need to keep monitoring this further, but what I think might be happening; After a period of idle time, it looks like the first activities from earlier staff members need to spin up a few things before they start moving again properly, causing this spike. Though the logs are referencing projects that no one has touched for weeks or in some cases years.

The most recent outage (the one I screenshotted the spike for) seems to give a number of errors that point to this (clipped to only interesting ones);

gitlab_error.log:

This is everything in the log

2026/08/12 06:43:12 [crit] 3945019#0: *1 connect() to unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket failed (2: No such file or directory) while connecting to upstream, client: 10.10.10.10, server: gitlab.example.com, request: "POST /api/v4/jobs/request HTTP/1.1", upstream: "http://unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket:/api/v4/jobs/request", host: "gitlab.example.com"
2026/08/12 06:43:12 [crit] 3945019#0: *1 connect() to unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket failed (2: No such file or directory) while connecting to upstream, client: 10.10.10.10, server: gitlab.example.com, request: "POST /api/v4/jobs/request HTTP/1.1", upstream: "http://unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket:/api/v4/jobs/request", host: "gitlab.example.com"
2026/08/12 06:43:12 [crit] 3945019#0: *1 connect() to unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket failed (2: No such file or directory) while connecting to upstream, client: 10.10.10.10, server: gitlab.example.com, request: "POST /api/v4/jobs/request HTTP/1.1", upstream: "http://unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket:/api/v4/jobs/request", host: "gitlab.example.com"
2026/08/12 06:43:12 [crit] 3945018#0: *5 connect() to unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket failed (2: No such file or directory) while connecting to upstream, client: 10.10.10.11, server: gitlab.example.com, request: "POST /api/v4/jobs/request HTTP/1.1", upstream: "http://unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket:/api/v4/jobs/request", host: "gitlab.example.com"
2026/08/12 06:43:12 [crit] 3945018#0: *5 connect() to unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket failed (2: No such file or directory) while connecting to upstream, client: 10.10.10.11, server: gitlab.example.com, request: "POST /api/v4/jobs/request HTTP/1.1", upstream: "http://unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket:/api/v4/jobs/request", host: "gitlab.example.com"
2026/08/12 06:43:12 [crit] 3945019#0: *1 connect() to unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket failed (2: No such file or directory) while connecting to upstream, client: 10.10.10.10, server: gitlab.example.com, request: "POST /api/v4/jobs/request HTTP/1.1", upstream: "http://unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket:/api/v4/jobs/request", host: "gitlab.example.com"

puma_stderr.log

=== puma startup: 2026-08-12 06:45:57 +0100 ===
DeclarativePolicy: large number of steps (54), falling back to static sort
DeclarativePolicy: large number of steps (54), falling back to static sort
DeclarativePolicy: large number of steps (53), falling back to static sort
DeclarativePolicy: large number of steps (54), falling back to static sort
DeclarativePolicy: large number of steps (54), falling back to static sort
DeclarativePolicy: large number of steps (54), falling back to static sort
DeclarativePolicy: large number of steps (53), falling back to static sort
DeclarativePolicy: large number of steps (53), falling back to static sort
# 86 lines of this

production_json.log

This is a snippet of one entry that gave a 401 status, and they seem to repeat regularly all the time, as in, 24 hours a day. There are a few of these within close proximity of this time frame, but everything after gives 200 with a couple 302’s. The weird thing is that a lot of the time it’s referencing projects that haven’t been touched by anyone for weeks or months. This is also confusing us and we’re not sure if it’s related or a different issue altogether.

{
    "method": "GET",
    "path": "/client-projects/client/client-project.git/info/refs",
    "format": "*/*",
    "controller": "Repositories::GitHttpController",
    "action": "info_refs",
    "status": 401,
    "time": "2026-08-12T06:44:44.722Z",
    "params": [
        {
            "key": "service",
            "value": "git-upload-pack"
        },
        {
            "key": "repository_path",
            "value": "client-projects/client/client-project.git"
        }
    ],
    "correlation_id": "01KZTBEPDRR8Y4B1R5R10QDXXE",
    "meta.caller_id": "Repositories::GitHttpController#info_refs",
    "meta.feature_category": "source_code_management",
    "repository_storage": "default",
    "remote_ip": "10.10.10.10",
    "ua": "git/2.50.1",
    "request_urgency": "default",
    "target_duration_s": 1,
    "db_count": 4,
    "db_write_count": 0,
    "db_cached_count": 0,
    "db_txn_count": 0,
    "db_replica_txn_count": 0,
    "db_primary_txn_count": 0,
    "db_replica_count": 0,
    "db_primary_count": 4,
    "db_replica_write_count": 0,
    "db_primary_write_count": 0,
    "db_replica_cached_count": 0,
    "db_primary_cached_count": 0,
    "db_replica_wal_count": 0,
    "db_primary_wal_count": 0,
    "db_replica_wal_cached_count": 0,
    "db_primary_wal_cached_count": 0,
    "db_replica_txn_max_duration_s": 0.0,
    "db_primary_txn_max_duration_s": 0.0,
    "db_replica_txn_duration_s": 0.0,
    "db_primary_txn_duration_s": 0.0,
    "db_replica_duration_s": 0.0,
    "db_primary_duration_s": 0.041,
    "db_main_txn_count": 0,
    "db_ci_txn_count": 0,
    "db_main_replica_txn_count": 0,
    "db_ci_replica_txn_count": 0,
    "db_main_count": 4,
    "db_ci_count": 0,
    "db_main_replica_count": 0,
    "db_ci_replica_count": 0,
    "db_main_write_count": 0,
    "db_ci_write_count": 0,
    "db_main_replica_write_count": 0,
    "db_ci_replica_write_count": 0,
    "db_main_cached_count": 0,
    "db_ci_cached_count": 0,
    "db_main_replica_cached_count": 0,
    "db_ci_replica_cached_count": 0,
    "db_main_wal_count": 0,
    "db_ci_wal_count": 0,
    "db_main_replica_wal_count": 0,
    "db_ci_replica_wal_count": 0,
    "db_main_wal_cached_count": 0,
    "db_ci_wal_cached_count": 0,
    "db_main_replica_wal_cached_count": 0,
    "db_ci_replica_wal_cached_count": 0,
    "db_main_txn_max_duration_s": 0.0,
    "db_ci_txn_max_duration_s": 0.0,
    "db_main_replica_txn_max_duration_s": 0.0,
    "db_ci_replica_txn_max_duration_s": 0.0,
    "db_main_txn_duration_s": 0.0,
    "db_ci_txn_duration_s": 0.0,
    "db_main_replica_txn_duration_s": 0.0,
    "db_ci_replica_txn_duration_s": 0.0,
    "db_main_duration_s": 0.041,
    "db_ci_duration_s": 0.0,
    "db_main_replica_duration_s": 0.0,
    "db_ci_replica_duration_s": 0.0,
    "path_traversal_check_duration_s": 0.000391,
    "cpu_s": 0.056876,
    "mem_objects": 10865,
    "mem_bytes": 4106944,
    "mem_mallocs": 2444,
    "mem_total_bytes": 4541544,
    "pid": 3945374,
    "worker_id": "puma_0",
    "rate_limiting_gates": [],
    "db_duration_s": 0.04086,
    "view_duration_s": 0.00077,
    "duration_s": 0.07473
}

gitlab-workhorse/current would have been interesting, but doesn’t go back far enough? There’s a number of executable files with strange names like @400000006a65d2ad38cccf44.s but nothing else. The config file just has the following

s209715200
n30
t86400
!gzip

Thanks for confirming. I asked because I sometimes see both on one VM competing for resources. My first GitLab in 2016 was exactly that, until I learned about separating them.

I’ve edited your post to obfuscate the domain in the logs.

It looks like something is restarting GitLab, rather than a resource constraint.

What the logs say

The NGINX error is neither a timeout nor a refused connection:

connect() to unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket failed (2: No such file or directory)

No such file or directory means the socket was gone, so Workhorse was not running. A saturated Workhorse would still own its socket and either accept slowly or refuse.

puma_stderr.log then shows a cold start with === puma startup: 2026-08-12 06:45:57 +0100 ===. Your NGINX errors begin at 06:43:12, so that is a restart of roughly three minutes, matching the one or two failed checks from Kuma.

The process IDs confirm it. NGINX workers 3945018 and 3945019, Puma worker 3945374, all allocated within moments of each other. The whole service tree came up together rather than one component recycling. The DeclarativePolicy lines are boot noise from that same start.

06:43 is also before the 08:00 start, so nobody was waiting on a cold cache.

Prime suspect: automatic package upgrades

On Ubuntu, apt-daily-upgrade.timer runs at 06:00 with a randomized delay of up to 60 minutes, so unattended upgrades land between 06:00 and 07:00 on a different minute every day. 06:43 sits in that window.

If gitlab-ee is in scope for unattended upgrades, the post-install step runs gitlab-ctl reconfigure and restarts every service. That explains both the outage and the CPU spike, since a Chef run plus a full Rails boot is expensive.

To confirm, on the GitLab VM:

grep -i gitlab /var/log/apt/history.log
systemctl list-timers apt-daily-upgrade.timer
journalctl --since "2026-08-12 06:30" --until "2026-08-12 07:00"

A Start-Date around 06:43 on 12 August with gitlab-ee in the upgrade list settles it.

The other things you spotted

The 401 is normal. Git over HTTP sends the first info/refs request without credentials, receives a 401, and retries with authentication. That is why it repeats around the clock and is immediately followed by 200s. The git/2.50.1 user agent and the runner IP tell you it is a CI job cloning, so if those repositories look untouched, check whether they have scheduled pipelines or are being mirrored.

Your Workhorse history is also still there. The @400000006a65d2ad38cccf44.s files are rotated logs from svlogd, not executables, gzipped by the processor in your config:

zcat /var/log/gitlab/gitlab-workhorse/@400000006a65d2ad38cccf44.s | less

The filename is a TAI64N timestamp, so tai64nlocal helps you find the file covering a given outage. Your config keeps 30 files rotated daily, so there is a month of history.

Next steps

Can you post the apt history output? If the upgrade is in there, the fix is to pin the package with apt-mark hold gitlab-ee and run upgrades manually.

We upgrade GitLab manually. It’s not part of our automatic upgrades. The one on the 12th completed before the outage occured;

Start-Date: 2026-08-12  06:41:40
Commandline: /usr/bin/unattended-upgrade
Upgrade: udev:amd64 (255.4-1ubuntu8.16, 255.4-1ubuntu8.17), libpam-systemd:amd64 (255.4-1ubuntu8.16, 255.4-1ubuntu8.17), libsystemd0:amd64 (255.4-1ubuntu8.16, 255.4-1ubuntu8.17), libnss-systemd:amd64 (255.4-1ubuntu8.16, 255.4-1ubuntu8.17), systemd:amd64 (255.4-1ubuntu8.16, 255.4-1ubuntu8.17), libudev1:amd64 (255.4-1ubuntu8.16, 255.4-1ubuntu8.17), systemd-dev:amd64 (255.4-1ubuntu8.16, 255.4-1ubuntu8.17), systemd-resolved:amd64 (255.4-1ubuntu8.16, 255.4-1ubuntu8.17), libsystemd-shared:amd64 (255.4-1ubuntu8.16, 255.4-1ubuntu8.17), systemd-sysv:amd64 (255.4-1ubuntu8.16, 255.4-1ubuntu8.17)
End-Date: 2026-08-12  06:42:00

Are one of these libraries something that causes GitLab to restart itself?

Yes, systemd and libsystemd0 are the ones that cause the OS to restart GitLab as a dependency.

Your apt run finished at 06:42:00 and the NGINX socket errors start at 06:43:12, about a minute later. needrestart runs after unattended upgrades, and when systemd or libsystemd0 is refreshed, its /etc/needrestart/restart.d/systemd-manager hook calls systemctl daemon-reexec and restarts units still linked against the old libraries. This behaviour on Ubuntu 24.04 is discussed on the ubuntu-devel list.

The reason it takes out all of GitLab at once is that a Linux package installation runs every component under runit, supervised by a single gitlab-runsvdir.service unit. Restart that one unit and you restart every GitLab service, which explains the missing Workhorse socket and why Puma needed until 06:45:57 to finish booting.

Confirm it on the VM:

journalctl -u gitlab-runsvdir.service --since "2026-08-12 06:40" --until "2026-08-12 06:50"
grep -iE "restart" /var/log/unattended-upgrades/unattended-upgrades.log

This is expected behavior: Ubuntu is doing what it is configured to do, and GitLab is coming back up as expected. It may need your evaluation on how you want to handle automatic service restarts.

One thing worth checking first

This explains 12 August nicely, but I would not call it solved just yet. Could you compare your other Kuma outages against the apt history?

grep "Start-Date" /var/log/apt/history.log

If they all land in that 06:00 to 07:00 window, this explains the root cause. If any happened during working hours, there is a second cause worth its own look, and the partially loading pages your users mentioned would more likely belong to that than to this restart.

This is great, thank you.

If any happened during working hours, there is a second cause worth its own look, and the partially loading pages your users mentioned would more likely belong to that than to this restart.

Some of them tie into apt, but some of them do not (the ones from 3am for example). We’ll monitor this more closely going forwards and will comment on this thread should it happen again outside of updates, especially since now we know better where and what to look for.

Thanks again for helping us solve this mystery.