← All posts

GitHub Actions on Blacksmith: test times and Docker cache costs

We moved GitHub Actions jobs for two TypeScript projects to Blacksmith: an Expo/React Native frontend and two Next.js apps built on Open Mercato. The jobs we'd been waiting 7–11 minutes for finished in roughly 3–8 minutes. Here's what those jobs actually do, what changed in the workflows, and why Docker cache storage needed a separate fix.

Time per CI job, before and afterMean job duration (minutes:seconds), 153 successful job executions. Each job kept its vCPU count. Database E2E also spans a SQL Server upgrade; details below.
A · Frontend checks
GitHub9:06
Blacksmith7:32
A · ERP integration
GitHub8:06
Blacksmith5:13
B · Checks + build
GitHub10:45
Blacksmith4:58
B · App E2E
GitHub7:08
Blacksmith3:22
B · Database E2E
GitHub10:04
Blacksmith4:50

What we were running

These are ordinary application pipelines: install dependencies, generate code, check types, run tests and, in some jobs, build the app. The projects are anonymized as A and B.

JobStackWhat's inside the jobRunner
A: frontend checksExpo 56, React Native 0.85, React 19, TypeScript, Bun, JestDependency install, Lingui translations, ESLint, typecheck and Jest. No production build or browser E2E.2 vCPU
A: ERP integrationNode 24, Yarn 4, Next.js 16, Open Mercato, Jest, PlaywrightInstall, code generation, typecheck, unit and ephemeral integration tests. Postgres 17/pgvector, Redis 7 and Meilisearch service containers.8 vCPU
B: checks + buildNode 24, Yarn 4, Next.js 16, Open Mercato, JestInstall, code generation, migrate/seed Postgres 17/pgvector, lint, typecheck, unit tests and a production Next.js build.2 vCPU
B: app E2ESame Next.js stack, PlaywrightBuild and start the production app against disposable Postgres, then test it through the running server.2 vCPU
B: database E2ESame stack, plus SQL ServerSeed a SQL Server container with the external ERP's schema, then run integration tests against the app and databases.2 vCPU

Open Mercato is the open-source business-app framework these Next.js apps use. In practice, that adds generated module code, database setup and integration tests to the usual JavaScript checks. The frontend is a separate Expo app; its install and test commands run through Bun.

What changed in CI

For the test jobs, we changed the runner label while keeping the vCPU count. A 2-vCPU Linux job became:

runs-on: blacksmith-2vcpu-ubuntu-2404

The scripts stayed in place: bun install --frozen-lockfile and the frontend checks, or yarn install --immutable, yarn generate, yarn typecheck, yarn test and the relevant integration/build commands. We left short routing jobs and deployment steps on their existing runners.

We sampled 153 successful first-attempt job executions: 15 before per job, then 15 after for project A and 16 for project B. Before runs are from September 7; after runs from September 9–10. The chart shows mean time from job start to finish, including setup and dependency installation, excluding queue time.

These were different commits close to the migration, not repeated builds of one frozen commit. Project B also changed a path filter. Its database E2E row includes a SQL Server upgrade from 2019 to 2025, so that row cannot tell us how much of the change came from the runner alone.

What a run costs

For 2-vCPU Linux runners, the rates we compared were $0.006/minute on GitHub and $0.004/minute on Blacksmith. Shorter jobs at a lower rate gave us an estimated 49–72% reduction in gross compute cost per sampled job.

For project B's three jobs together, that was about $0.178 → $0.053 per run of the set. These jobs run in parallel, so adding their costs makes sense; adding their durations would not tell us how long a developer waits.

The estimate uses GitHub's per-job whole-minute rounding and Blacksmith's elapsed time × rate. If we round Blacksmith jobs up to whole minutes too, the saving for that three-job set is about 67%, rather than 70%. Rates checked October 2: GitHub, Blacksmith.

The full bill, including storage and credits

The monthly total covers the whole organization, including jobs outside those five samples.

Net metered CI charges, USDAugustSeptember
GitHub, including compute, storage, registry usage and allowances$640.41$378.91
Blacksmith compute$0.00$243.47
Blacksmith Docker cache storage$0.00$205.54
Blacksmith invoice credits$0.00−$24.00
Combined total$640.41$803.92

September cost 25.5% more. It also had a different volume and mix of work, and the rollout happened during the month. That isn't a like-for-like price comparison. It does show why a cheaper test run isn't enough to claim a lower total bill.

Both $12 Blacksmith invoice credits are included. GitHub figures are net usage-ledger charges; base subscriptions, invoice taxes and account-level adjustments are outside this comparison.

Docker builds were a separate change

The five timing results above are CI test/check jobs. They don't measure the Docker layer-cache change.

The Next.js production images use multistage node:24-alpine Dockerfiles: install with Yarn, generate framework code, build Next.js, then copy the app into a runtime stage with production dependencies. For those image jobs, we switched to Blacksmith's Buildx setup and build/push actions, with a fixed cache key per image:

- uses: useblacksmith/setup-docker-builder@v2
  with:
    cache-key: app-image

- uses: useblacksmith/build-push-action@v2
  with:
    context: .
    push: true
    tags: ghcr.io/your-org/your-app:your-tag

The key stays the same between commits. BuildKit decides which layers are reusable. The sticky disk stores its build cache between jobs, replacing the external cache-from/cache-to export we used before.

That storage costs $0.50/GB-month. Our October 2 inventory had paid caches of approximately 143 GB, 96 GB and 37 GB. At those sizes all month, that's about $138. September's actual charge was $205.54; today's sizes don't reconstruct the older caches or their growth.

The default cleanup is based on how long layers have gone unused. Our live builder used a 72-hour policy, overriding the action's eight-day default. Recently used layers had no configured size budget. See Blacksmith's cache documentation.

What we changed to reduce storage

First, we removed .mercato/next/cache/turbopack in the same Dockerfile RUN as the production build. It's compiler state, so we don't need to ship it with the running app. Deleting it in a later layer would leave the bytes in the earlier one.

Second, we tried a 30 GB retained-cache budget on the largest cache. After a successful image build, before Blacksmith commits the disk:

docker buildx du
docker buildx prune --force --max-used-space 30000000000
docker buildx du

That's 30 billion bytes. Docker's 30gb suffix means 30 GiB. This command targets the selected builder's reclaimable cache, not a hard limit on the whole disk. Check your BuildKit version supports the option; ours was 0.29.6-blacksmith. Prune reference.

First live trialResult
Cache after building, before pruning156 GB
Cache after pruning27.41 GB
Data reclaimed128.6 GB
Pruning time94 seconds
Image build and push time4m48s

The image, tests and preview deployment passed. The dashboard had not shown a corresponding storage reduction when we checked, so this is a successful cleanup, not a measured bill saving yet. The PRs are still in review. All writers of the trial cache use one shared workflow; that matters because another writer using the old policy can save a larger cache again.

What the limit might have saved in September

We modeled 30 GB per key across the nine cache keys found in September's workflow history, including retired ones. Using approximate introduction dates from migration PRs and repository creation gives 162 cache-days. Retaining all nine for the full September 8–30 window gives 207 cache-days.

At $0.50/GB-month, a 30 GB cache costs $0.50 per September day:

September scenarioStorageCombined CI charges
Actual usage$205.54$803.92
30 GB per cache, approximate introduction dates$81.00$679.38
30 GB per cache, all nine retained for 23 days$103.50$701.88

Those estimates keep both credits and the other charges unchanged. They assume the limit reduces billed storage. They don't include the cost of extra rebuilding.

Put an illustrative $10–30 aside for that extra compute, and the modeled saving is $72–115 for September. For scale, 1,000 builds adding five minutes each on a $0.004/minute runner cost $20. We haven't measured that penalty yet; the first 94-second prune isn't a prediction for every later build.

For our setup, the shorter test jobs are a useful result already. The next check is whether normal builds stay quick with less retained cache, and whether that smaller cache actually reduces the storage charge. If you're making the same move, record those two things from day one.

Bring the napkin sketch.

Thirty minutes with a full-stack engineer, plus a short written report that's yours to keep.

30 minGoogle Calendarno sales deck

Prefer to write? Send us a message →