Hi there,
So, reviewing the logs you pointed me to (thank you!), I’m also starting to agree that there’s a constraint somewhere, and now I’m trying to figure out the where and why part.
Further information on our setup; we don’t have Gitlab runners on the Gitlab VM itself. A separate VM is in place with a single GitLab Runner in a container to receive the jobs. When a job gets fed through the pipeline, that runner receives it and spins up a fresh, disposable “runner VM” to execute the job (tags determine whether it’s a general purpose VM or a high powered VM for larger projects). have CI runners that get spun up via a ‘disposable’ VM. (Effectively this)
I don’t have logs for all of the downtime events, so I’ll need to keep monitoring this further, but what I think might be happening; After a period of idle time, it looks like the first activities from earlier staff members need to spin up a few things before they start moving again properly, causing this spike. Though the logs are referencing projects that no one has touched for weeks or in some cases years.
The most recent outage (the one I screenshotted the spike for) seems to give a number of errors that point to this (clipped to only interesting ones);
gitlab_error.log:
This is everything in the log
2026/08/12 06:43:12 [crit] 3945019#0: *1 connect() to unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket failed (2: No such file or directory) while connecting to upstream, client: 10.10.10.10, server: gitlab.example.com, request: "POST /api/v4/jobs/request HTTP/1.1", upstream: "http://unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket:/api/v4/jobs/request", host: "gitlab.example.com"
2026/08/12 06:43:12 [crit] 3945019#0: *1 connect() to unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket failed (2: No such file or directory) while connecting to upstream, client: 10.10.10.10, server: gitlab.example.com, request: "POST /api/v4/jobs/request HTTP/1.1", upstream: "http://unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket:/api/v4/jobs/request", host: "gitlab.example.com"
2026/08/12 06:43:12 [crit] 3945019#0: *1 connect() to unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket failed (2: No such file or directory) while connecting to upstream, client: 10.10.10.10, server: gitlab.example.com, request: "POST /api/v4/jobs/request HTTP/1.1", upstream: "http://unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket:/api/v4/jobs/request", host: "gitlab.example.com"
2026/08/12 06:43:12 [crit] 3945018#0: *5 connect() to unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket failed (2: No such file or directory) while connecting to upstream, client: 10.10.10.11, server: gitlab.example.com, request: "POST /api/v4/jobs/request HTTP/1.1", upstream: "http://unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket:/api/v4/jobs/request", host: "gitlab.example.com"
2026/08/12 06:43:12 [crit] 3945018#0: *5 connect() to unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket failed (2: No such file or directory) while connecting to upstream, client: 10.10.10.11, server: gitlab.example.com, request: "POST /api/v4/jobs/request HTTP/1.1", upstream: "http://unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket:/api/v4/jobs/request", host: "gitlab.example.com"
2026/08/12 06:43:12 [crit] 3945019#0: *1 connect() to unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket failed (2: No such file or directory) while connecting to upstream, client: 10.10.10.10, server: gitlab.example.com, request: "POST /api/v4/jobs/request HTTP/1.1", upstream: "http://unix:/var/opt/gitlab/gitlab-workhorse/sockets/socket:/api/v4/jobs/request", host: "gitlab.example.com"
puma_stderr.log
=== puma startup: 2026-08-12 06:45:57 +0100 ===
DeclarativePolicy: large number of steps (54), falling back to static sort
DeclarativePolicy: large number of steps (54), falling back to static sort
DeclarativePolicy: large number of steps (53), falling back to static sort
DeclarativePolicy: large number of steps (54), falling back to static sort
DeclarativePolicy: large number of steps (54), falling back to static sort
DeclarativePolicy: large number of steps (54), falling back to static sort
DeclarativePolicy: large number of steps (53), falling back to static sort
DeclarativePolicy: large number of steps (53), falling back to static sort
# 86 lines of this
production_json.log
This is a snippet of one entry that gave a 401 status, and they seem to repeat regularly all the time, as in, 24 hours a day. There are a few of these within close proximity of this time frame, but everything after gives 200 with a couple 302’s. The weird thing is that a lot of the time it’s referencing projects that haven’t been touched by anyone for weeks or months. This is also confusing us and we’re not sure if it’s related or a different issue altogether.
{
"method": "GET",
"path": "/client-projects/client/client-project.git/info/refs",
"format": "*/*",
"controller": "Repositories::GitHttpController",
"action": "info_refs",
"status": 401,
"time": "2026-08-12T06:44:44.722Z",
"params": [
{
"key": "service",
"value": "git-upload-pack"
},
{
"key": "repository_path",
"value": "client-projects/client/client-project.git"
}
],
"correlation_id": "01KZTBEPDRR8Y4B1R5R10QDXXE",
"meta.caller_id": "Repositories::GitHttpController#info_refs",
"meta.feature_category": "source_code_management",
"repository_storage": "default",
"remote_ip": "10.10.10.10",
"ua": "git/2.50.1",
"request_urgency": "default",
"target_duration_s": 1,
"db_count": 4,
"db_write_count": 0,
"db_cached_count": 0,
"db_txn_count": 0,
"db_replica_txn_count": 0,
"db_primary_txn_count": 0,
"db_replica_count": 0,
"db_primary_count": 4,
"db_replica_write_count": 0,
"db_primary_write_count": 0,
"db_replica_cached_count": 0,
"db_primary_cached_count": 0,
"db_replica_wal_count": 0,
"db_primary_wal_count": 0,
"db_replica_wal_cached_count": 0,
"db_primary_wal_cached_count": 0,
"db_replica_txn_max_duration_s": 0.0,
"db_primary_txn_max_duration_s": 0.0,
"db_replica_txn_duration_s": 0.0,
"db_primary_txn_duration_s": 0.0,
"db_replica_duration_s": 0.0,
"db_primary_duration_s": 0.041,
"db_main_txn_count": 0,
"db_ci_txn_count": 0,
"db_main_replica_txn_count": 0,
"db_ci_replica_txn_count": 0,
"db_main_count": 4,
"db_ci_count": 0,
"db_main_replica_count": 0,
"db_ci_replica_count": 0,
"db_main_write_count": 0,
"db_ci_write_count": 0,
"db_main_replica_write_count": 0,
"db_ci_replica_write_count": 0,
"db_main_cached_count": 0,
"db_ci_cached_count": 0,
"db_main_replica_cached_count": 0,
"db_ci_replica_cached_count": 0,
"db_main_wal_count": 0,
"db_ci_wal_count": 0,
"db_main_replica_wal_count": 0,
"db_ci_replica_wal_count": 0,
"db_main_wal_cached_count": 0,
"db_ci_wal_cached_count": 0,
"db_main_replica_wal_cached_count": 0,
"db_ci_replica_wal_cached_count": 0,
"db_main_txn_max_duration_s": 0.0,
"db_ci_txn_max_duration_s": 0.0,
"db_main_replica_txn_max_duration_s": 0.0,
"db_ci_replica_txn_max_duration_s": 0.0,
"db_main_txn_duration_s": 0.0,
"db_ci_txn_duration_s": 0.0,
"db_main_replica_txn_duration_s": 0.0,
"db_ci_replica_txn_duration_s": 0.0,
"db_main_duration_s": 0.041,
"db_ci_duration_s": 0.0,
"db_main_replica_duration_s": 0.0,
"db_ci_replica_duration_s": 0.0,
"path_traversal_check_duration_s": 0.000391,
"cpu_s": 0.056876,
"mem_objects": 10865,
"mem_bytes": 4106944,
"mem_mallocs": 2444,
"mem_total_bytes": 4541544,
"pid": 3945374,
"worker_id": "puma_0",
"rate_limiting_gates": [],
"db_duration_s": 0.04086,
"view_duration_s": 0.00077,
"duration_s": 0.07473
}
gitlab-workhorse/current would have been interesting, but doesn’t go back far enough? There’s a number of executable files with strange names like @400000006a65d2ad38cccf44.s but nothing else. The config file just has the following
s209715200
n30
t86400
!gzip