How I build Nest.js services
Close to a decade and dozens of production services in: the project structure, dependency injection rules, test harness, consumers, migrations, auth and monitoring I start every Nest.js project with, plus the skills that let an agent verify its own work.
Nine of every ten backend projects I’ve shipped over the last decade have been Nest.js. GraphQL APIs, REST APIs, microservices, RabbitMQ consumers, websocket handlers, CLIs, cron jobs: dozens of them, running in production, for employers and for clients. Somewhere along the way the framework stopped being the interesting part, and the conventions I’d built up around it started to be.
What follows is the set of conventions that survived. It’s what I’d set up on day one of a new project, assuming Apollo GraphQL and TypeORM on Postgres, with RabbitMQ for background work and Datadog watching it all, because that’s what most of my projects look like.
What Nest.js is, and why it stuck
Nest is a TypeScript framework for building Node.js servers. Kamil Myśliwiec started it in early 2017, borrowing Angular’s architecture (modules, providers, decorators and a dependency injection container) and running it on top of Express, with a Fastify adapter as the alternative. It’s MIT licensed, has about 76,000 stars on GitHub, and @nestjs/core was downloaded a little over ten million times from npm last week. Version 12 shipped at the end of August. Everything in this post is on 11, which is where I’d expect most production codebases to sit for a while yet.
The HTTP layer is the least interesting part. What Nest actually gives you is an opinion about where code goes and a container that wires it together. Every class that does work is a provider, providers live in modules, modules declare what they export, and at boot the container resolves the whole graph and hands each class its dependencies. First-party packages cover most of the rest of a backend: @nestjs/graphql with an Apollo driver (federation included), @nestjs/typeorm, @nestjs/config, @nestjs/schedule, @nestjs/testing, websocket gateways and a microservices transport layer. The community fills the gaps, and nest-commander for CLIs is the one I use on every project.
That opinion is the reason I’ve stayed. On a team, most of a framework’s value is that the next person can guess where things live, and Nest gets you a long way there before you’ve written a single convention of your own. The next person is now as likely to be an agent as a human, which makes guessability worth more than it used to be.
It isn’t free. It’s decorators all the way down, the container fails at boot rather than at compile time when you forget to export a provider, and major-version upgrades drag a long tail of dependencies behind them, a gripe I’ve made before. I’ve also admitted to being a bit bored of it. None of that has outweighed being able to open any Nest codebase I’ve touched in the last decade and find my way around it in minutes.
Where this comes from
Thoughts on GraphQL after 2 years covers one stretch of that decade: splitting a GraphQL service into federated subgraphs, each its own Nest application. Most of what follows was learned the same way, one production system at a time.
Two recent codebases are the reference points for this post, because they’re the most complete expression of how I work now. One is a federated GraphQL subgraph with a fleet of RabbitMQ consumers, from my day job. The other is a CRM monorepo I build and run for a client. Both are proprietary, so every sample below is rewritten against a made-up domain of articles, authors and comments. The shapes are real. The names aren’t.
Project structure
Domain modules, and GraphQL in its own
Modules are organised by domain. An articles module owns the article entities, repositories and services, a comments module owns comments, and none of them know GraphQL exists. The API is its own graphql module, which imports the domain modules and translates between them and the schema.
src/
main.ts # HTTP entrypoint
cli.ts # nest-commander entrypoint
tracer.ts # dd-trace, imported before anything else
consumers/
bootstrap.ts
article-published/
main.ts # one entrypoint per consumer
config.json
article-published.consumer.ts
article-published.consumer.provider.ts
modules/
app/
articles/ # a domain module
articles.module.ts
entities/
repositories/
services/
comments/
messenger/
observability/
cli/
cli.module.ts
commands/
graphql/ # the API layer
graphql.module.ts
objects/
inputs/
resolvers/
loaders/
filters/
The CRM is where I got this wrong. Early on I mapped GraphQL straight onto the TypeORM entities, one class carrying both sets of decorators:
// Don't: one class, two jobs.
@ObjectType()
@Entity({ name: 'articles' })
export class Article {
@Field(() => ID)
@PrimaryGeneratedColumn('uuid')
public readonly id!: string;
@Field(() => String)
@Column()
public title!: string;
@Field(() => Author)
@ManyToOne(() => Author)
@JoinColumn({ name: 'author_id' })
public author!: Author;
}
It’s genuinely quicker for the first dozen entities, which is exactly why it’s tempting. It gets worse with every entity after that.
The database schema becomes the API contract. Renaming a column is now a breaking change for every client, or it’s a column whose name stops matching its meaning for the rest of the codebase’s life.
Relations lie about nullability. author is only populated if someone remembered to load the relation, but the schema promises it’s always there. The fix is a field resolver that shadows the property, at which point the entity is declaring a field it doesn’t actually serve.
GraphQL leaks downward. Enum registrations end up in entity files, and anything that imports an entity, like a queue consumer or a CLI command, drags @nestjs/graphql in with it. And every column that should never leave the database is one misplaced @Field() away from being in the schema.
The version I’d build today keeps them apart. The object type is its own class in the GraphQL module:
// graphql/objects/article.object.ts
@ObjectType({ description: 'An article, published or draft.' })
@Directive('@key(fields: "id")')
export class Article {
@Field(() => ID)
public readonly id!: string;
@Field(() => String)
public readonly title!: string;
@Field(() => ArticleStatus)
public readonly status!: ArticleStatus;
@Field(() => Date, { nullable: true })
public readonly publishedAt!: Date | null;
}
Relations stop being properties that may or may not be loaded and become field resolvers backed by a loader. The separation also costs less than it looks, because TypeScript is structural: a resolver typed to return the object type can return the entity directly while their shapes agree. When they stop agreeing, the compiler tells you, rather than a client.
The flow through the layer is GraphQL input, resolver, plain DTO, service. Services and repositories take plain TypeScript types, never an @InputType() class, so nothing beneath the resolver imports @nestjs/graphql. That pays off later, when the same domain modules are booted by a queue consumer or a CLI command with no reason to build a schema.
A folder per service
Every service with any weight gets its own folder, with up to five files:
services/article-search/
article-search.service.ts # the @Injectable() class, nothing else
article-search.service.types.ts # ArticleSearchServiceOptions, internal types
article-search.service.constants.ts # module-scope constants
article-search.service.support.ts # pure functions
article-search.service.provider.ts # FactoryProvider, only if setup is needed
The service file holds the class and its imports. Anything that doesn’t touch this moves to support.ts as a pure function and gets unit-tested there, which turns out to be where nearly all of the unit tests live. Constants live apart so the service, its support functions and the tests share one definition. The options type lives apart so the provider and the specs can import it without importing the class. The provider only exists if the service needs setup.
It’s fussy for a thirty-line service, so the rule is to skip the files you don’t need and never create empty placeholders. It earns its keep on the service that’s grown to six hundred lines, where “what does this depend on, and what does it do on its own?” have become hard questions, and the folder answers both before you open a file.
Dependency injection
Three rules, and they’re really one rule: a service receives fully formed collaborators and does work with them. How those collaborators came to exist is a provider’s problem.
Services never see config
ConfigService never gets injected into a service. A FactoryProvider reads config, pulls out the handful of values the service needs, and passes them in as a typed options object, always the first constructor argument:
// article-search.service.types.ts
export interface ArticleSearchServiceOptions {
/** Index that published articles are written to. */
readonly indexName: string;
/** Maximum number of articles sent per indexing request. */
readonly batchSize: number;
}
// article-search.service.provider.ts
export const ArticleSearchServiceProvider: FactoryProvider<ArticleSearchService> = {
provide: ArticleSearchService,
inject: [ConfigService, Logger, ArticleRepository, SearchClient],
useFactory: (
config: ConfigService,
logger: Logger,
articles: ArticleRepository,
search: SearchClient,
) =>
new ArticleSearchService(
{
indexName: config.getOrThrow<string>('search.articleIndex'),
batchSize: config.get<number>('search.batchSize', 100),
},
logger,
articles,
search,
),
};
// article-search.service.ts
@Injectable()
export class ArticleSearchService {
constructor(
private readonly options: ArticleSearchServiceOptions,
private readonly logger: Logger,
private readonly articles: ArticleRepository,
private readonly search: SearchClient,
) {}
}
A service holding ConfigService can read any configuration value in the application, and given enough time it will. The options type is the opposite: an exhaustive, documented list of what the service depends on, visible without reading its body. Tests construct the service with a literal and there’s nothing to mock. And getOrThrow in a factory fails at boot, naming the missing variable, instead of on the first request that happens to need it. The module registers ArticleSearchServiceProvider, not the bare class.
Nothing initialises itself
If a dependency needs asynchronous setup, like a TCP connection or a broker channel, that setup happens in an async factory and the service receives the finished product:
export const MessengerServiceProvider: FactoryProvider<MessengerService> = {
provide: MessengerService,
inject: [ConfigService, Logger],
useFactory: async (config: ConfigService, logger: Logger) => {
const connection = await connect({
hostname: config.getOrThrow('amqp.host'),
port: config.getOrThrow<number>('amqp.port'),
username: config.getOrThrow('amqp.username'),
password: config.getOrThrow('amqp.password'),
vhost: config.getOrThrow('amqp.vhost'),
heartbeat: 10,
});
const channel = await connection.createChannel();
await channel.prefetch(config.get<number>('amqp.prefetch', 1));
return new MessengerService(connection, channel, logger);
},
};
Nest awaits an async factory before it constructs anything that depends on it. So the service never sees a half-connected client, never needs an “are we connected yet?” branch in every method, and a broker the app can’t reach is a failed boot instead of a failed first message. The alternative, calling connect() in onModuleInit or lazily on first use, gets all three wrong. The service also stays easy to test: hand it a fake channel.
The same service tracks a healthy flag, flipped by the connection’s error and close events, which the readiness probe reads. More on that under monitoring.
Nothing constructs its own clients
The same goes for HTTP clients. The provider builds the axios instance with its base URL, auth headers and timeout, and the service takes an AxiosInstance:
export const ProfileServiceProvider: FactoryProvider<ProfileService> = {
provide: ProfileService,
inject: [ConfigService],
useFactory: (config: ConfigService) =>
new ProfileService(
axios.create({
baseURL: config.getOrThrow('profiles.url'),
timeout: 5_000,
headers: {
'x-requested-with': 'articles-service',
'x-internal-token': config.getOrThrow('profiles.internalToken'),
},
}),
),
};
@Injectable()
export class ProfileService {
constructor(private readonly client: AxiosInstance) {}
}
A service that calls axios.create() in its own constructor has hard-coded where it talks to, how it authenticates and how long it’s willing to wait. Pass it a prepared instance and an integration test can pass a different one pointed at WireMock. No module overrides, no environment juggling, and the code under test is the code that ships.
The GraphQL layer
Loaders are singletons with the cache off
Every field resolver that fetches anything goes through a DataLoader. That part is standard. The less standard part is scope.
Nest’s request scope bubbles up the injection chain: inject one request-scoped provider and every class that depends on it becomes request-scoped too, rebuilt for every request. The usual advice is to make loaders request-scoped so their cache can’t leak between users. I do the opposite. Loaders are ordinary singletons with the cache disabled.
DataLoader batches the keys requested within the same tick whether it caches or not, so a singleton loader with cache: false is just a batching function, and under load it will happily batch across concurrent requests. In my experience the per-request memoisation you give up has rarely been worth the request-scoped DI it costs. The one rule it adds: anything user-specific goes into the key, never into the loader.
@Injectable()
export class CommentCountForArticleLoader extends ObservableDataLoader<string, number> {
constructor(metrics: Metrics, comments: CommentRepository) {
super(
async (articleIds) => {
const counts = await comments.getCountsForArticles(articleIds);
// Results must come back in the same order as the keys.
return articleIds.map((id) => counts.get(id) ?? 0);
},
{ metricsAgent: metrics, cache: false },
);
}
}
@ResolveField(() => Int)
public commentCount(@Parent() { id }: Pick<Article, 'id'>): Promise<number> {
return this.commentCountForArticleLoader.load(id);
}
The rest of the rules are about correctness. One repository call per batch, never a loop. Results come back in key order, so every loader ends in keys.map(...). A missing entity that the schema says is required becomes a NotFoundException in that key’s slot, which rejects just that one load() rather than the whole batch. The repository method underneath takes an array and returns a Map, which I’ve written about before. ObservableDataLoader is a thin subclass that adds a trace span and batch-size metrics, and it comes back under monitoring.
Lists and counts share one query
Paginated list queries all follow one shape. A repository method builds the filtered query and returns it unpaginated and unsorted, the list resolver adds ordering and pagination, and a sibling count resolver calls getCount() on the same builder:
@Query(() => [Article])
public articles(
@Args('input') input: GetArticlesInput,
@Args('offset', { type: () => Int, defaultValue: 0 }) offset: number,
@Args('limit', { type: () => Int, defaultValue: 100 }) limit: number,
): Promise<readonly Article[]> {
return this.articleRepository
.buildSearchQuery(input)
// orderBy and order were whitelisted by input validation.
.orderBy(`article.${input.orderBy}`, input.order)
.skip(offset)
.take(limit)
.getMany();
}
@Query(() => Int)
public articlesCount(@Args('input') input: GetArticlesInput): Promise<number> {
return this.articleRepository.buildSearchQuery(input).getCount();
}
The filters exist once, so the list and the count can’t disagree, and an E2E test asserting they agree is the cheap way to keep it that way. orderBy gets interpolated into SQL, so the input’s validation schema whitelists it against known columns first; a free-form string never reaches .orderBy(). For plain reads like this the resolver can go straight to the repository. Anything that changes state goes through a service.
Offset pagination is fine for a CRM’s hundreds-to-thousands of rows, and much simpler for the frontend. The subgraph uses cursor connections, because some of its lists are much larger.
Testing
The shape I aim for puts most of the weight in the middle. Unit tests cover logic that lives entirely inside a function. Integration tests cover one service or repository against real infrastructure, including the HTTP services it calls. E2E tests cover the external surface, meaning GraphQL, REST and queue consumers, through the whole application.
Unit tests are for logic
A unit test that mocks a repository to assert the service called the repository mostly proves the mock was configured. Those tests are cheap to write and expensive to keep, and they break on every refactor while rarely catching a real bug. So unit tests are reserved for functions whose behaviour is entirely their implementation: decisions, calculations, parsing, formatting. In practice that means the .support.ts files, which is part of why they exist.
// article.service.support.ts
export function readingTimeMinutes(html: string, wordsPerMinute = 230): number {
const words = html.replace(/<[^>]+>/g, ' ').split(/\s+/).filter(Boolean).length;
return Math.max(1, Math.ceil(words / wordsPerMinute));
}
describe('readingTimeMinutes', () => {
it('never reports less than a minute', () => {
expect(readingTimeMinutes('<p>Hi</p>')).toBe(1);
});
it('ignores markup when counting words', () => {
const html = `<p>${'word '.repeat(460)}</p><img src="cover.png" />`;
expect(readingTimeMinutes(html)).toBe(2);
});
});
No mocks, no setup, and they run in milliseconds.
Integration tests: one layer, real infrastructure
Integration tests don’t boot Nest at all. Each suite creates a standalone TypeORM DataSource against a real Postgres, using the same entity globs as the app, and constructs the thing under test by hand:
export function createDataSource() {
return new DataSource({
type: 'postgres',
host: process.env.DB_HOST,
port: Number(process.env.DB_PORT),
database: process.env.DB_NAME,
username: process.env.DB_USERNAME,
password: process.env.DB_PASSWORD,
entities: [resolve(__dirname, '../../src/**/entities/*.entity.ts')],
synchronize: false,
});
}
describe('ArticleRepository', () => {
let dataSource: DataSource;
let articles: ArticleRepository;
beforeAll(async () => {
dataSource = createDataSource();
await dataSource.initialize();
registerDataSource(dataSource);
articles = new ArticleRepository(dataSource);
});
describe('buildSearchQuery()', () => {
it('excludes archived articles by default', async () => {
const authorId = await seedAuthor(dataSource);
const live = await seedArticle(dataSource, { authorId });
await seedArticle(dataSource, { authorId, archived: true });
const query = articles.buildSearchQuery({ authorIds: [authorId] });
expect((await query.getMany()).map((a) => a.id)).toEqual([live]);
expect(await query.getCount()).toBe(1);
});
});
});
Test data goes in through raw SQL with faker defaults rather than through the entities. I like that it keeps each test’s arrangement explicit, and that an entity mapping bug can’t quietly arrange the very data it’s about to be tested against.
Cleanup is automatic. registerDataSource hands the connection to a global afterAll, which truncates every table derived from the entity metadata except a short, hand-kept list of lookup tables that migrations seed:
const SEEDED_TABLES = new Set(['article_statuses', 'comment_flags', 'regions']);
export async function resetDatabase(dataSource: DataSource) {
const tables = dataSource.entityMetadatas
.filter((entity) => entity.tableType !== 'view')
.map((entity) => entity.tableName)
.filter((table) => !SEEDED_TABLES.has(table));
if (tables.length === 0) return;
await dataSource.query(`TRUNCATE TABLE ${tables.map((t) => `"${t}"`).join(', ')}`);
}
A new entity gets truncated between suites with no change to the helper. A new lookup table needs adding to SEEDED_TABLES in the same change as its migration, which is exactly the kind of rule that ends up in a skill. One wrinkle: if a seeded table holds a foreign key into a transactional one, Postgres refuses the truncate outright, and CASCADE would take the seed data with it. The fix is to drop those few cross-boundary constraints, truncate, and put them back.
The value of this layer is focus. A repository method with five filters gets a test per filter, one for empty input and one with the filters combined, each running a real query through a real planner. The complex SQL I now write far more of, for reasons covered in the 2026 stack post, is only safe to write because this layer exists.
External services: WireMock and retries
The same layer covers services that call other services over HTTP, with WireMock standing in for the other end. It runs as a container beside Postgres, and tests program it through its admin API with a wrapper that’s barely worth showing:
export class WireMockClient {
constructor(private readonly baseUrl: string) {}
public async createMapping(mapping: StubMapping): Promise<void> {
const response = await fetch(`${this.baseUrl}/__admin/mappings`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify(mapping),
});
if (!response.ok) {
throw new Error(`WireMock rejected mapping: ${response.status} ${await response.text()}`);
}
}
public async resetMappings(): Promise<void> {
await fetch(`${this.baseUrl}/__admin/mappings/reset`, { method: 'POST' });
}
}
Because the provider owns the axios instance, the test builds its own, points it at WireMock, and constructs the real service with it:
describe('ProfileService', () => {
const wiremock = new WireMockClient(process.env.WIREMOCK_URL!);
let client: AxiosInstance;
let profiles: ProfileService;
beforeEach(async () => {
client = axios.create({
baseURL: process.env.WIREMOCK_URL,
headers: { 'x-internal-token': 'test-token' },
});
profiles = new ProfileService(client);
await wiremock.resetMappings();
});
it('sends the new display name, authenticated', async () => {
await wiremock.createMapping({
request: {
method: 'PATCH',
urlPath: '/internal/profiles/abc',
headers: { 'x-internal-token': { equalTo: 'test-token' } },
bodyPatterns: [{ equalToJson: JSON.stringify({ displayName: 'Marty' }) }],
},
response: {
status: 200,
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ id: 'abc', displayName: 'Marty' }),
},
});
await expect(profiles.updateDisplayName('abc', 'Marty')).resolves.toEqual({
id: 'abc',
displayName: 'Marty',
});
});
it('retries a 503 three times, then gives up', async () => {
jest.spyOn(client, 'patch');
await wiremock.createMapping({
request: { method: 'PATCH', urlPath: '/internal/profiles/abc' },
response: { status: 503 },
});
await expect(profiles.updateDisplayName('abc', 'Marty')).rejects.toThrow();
expect(client.patch).toHaveBeenCalledTimes(3);
}, 10_000);
it('does not retry a 404', async () => {
jest.spyOn(client, 'patch');
await wiremock.createMapping({
request: { method: 'PATCH', urlPath: '/internal/profiles/abc' },
response: { status: 404 },
});
await expect(profiles.updateDisplayName('abc', 'Marty')).rejects.toThrow();
expect(client.patch).toHaveBeenCalledTimes(1);
});
});
This is what makes WireMock worth the container. It isn’t limited to status codes: it can delay a response past your timeout, or fail at the socket with faults like CONNECTION_RESET_BY_PEER, EMPTY_RESPONSE and MALFORMED_RESPONSE_CHUNK. A mocked axios tells you the service called patch. It doesn’t tell you whether a reset connection gets classified as transient, whether your headers survive axios’s merging, or whether the error a real 503 produces has the shape your catch block assumes.
Most of those tests exist to pin down retries, which I’d rate among the best value-per-line code in any service that talks over a network. retry-as-promised does the work, and the service decides what counts as transient:
class TransientProfileError extends Error {}
@Injectable()
export class ProfileService {
constructor(private readonly client: AxiosInstance) {}
/**
* Setting a display name is idempotent, so a transient failure is safe to
* retry: sending it twice leaves the profile as sending it once would.
*/
public async updateDisplayName(id: string, displayName: string): Promise<Profile> {
return retryAsPromised(
async () => {
try {
const { data } = await this.client.patch<Profile>(`internal/profiles/${id}`, {
displayName,
});
return data;
} catch (error) {
// No response at all, or a 5xx, is worth another attempt. A 4xx won't
// improve with repetition, so it fails fast.
if (axios.isAxiosError(error)) {
const status = error.response?.status;
if (status === undefined || status >= 500) {
throw new TransientProfileError(error.message, { cause: error });
}
}
throw error;
}
},
{ max: 3, backoffBase: 1_000, backoffExponent: 2, match: [TransientProfileError] },
);
}
}
The condition that isn’t negotiable is idempotency. Setting a display name twice leaves the profile exactly as setting it once does, so retrying is safe. Creating an invoice twice is not, and a retry wrapped around a non-idempotent call is a duplicate-charge incident waiting for a flaky network. Those calls either carry an idempotency key the other side honours, or they don’t get retried.
E2E: the surface area
E2E tests boot the whole AppModule through @nestjs/testing and talk to it the way clients do: GraphQL and REST through supertest, messages through the real broker. Everything the application owns is real, down to Postgres. Everything it doesn’t own is replaced at the provider boundary:
export function buildAppModuleBuilder() {
return Test.createTestingModule({ imports: [AppModule] })
.overrideProvider(MailService)
.useValue(MailServiceMock.getDefault())
.overrideProvider(StorageService)
.useValue(StorageServiceMock.getDefault())
.overrideProvider(ProfileService)
.useValue(ProfileServiceMock.getDefault())
.overrideGuard(GraphAuthGuard)
.useValue({
canActivate: (context: ExecutionContext) => {
GqlExecutionContext.create(context).getContext().req.user = stubPrincipal();
return true;
},
});
}
That override list only works because of the DI rules. A service that built its own mail client couldn’t be swapped out without patching the module that declares it. External HTTP behaviour is already covered at the integration layer, so here a provider override is the cheaper stand-in. The subgraph points some of its E2E suites at WireMock instead, and either works. The auth guard is replaced with one that attaches a stub principal, an admin by default; specs that care about a particular role override it again, and per-permission behaviour gets unit tests against the guard itself.
Specs come in two flavours. Some isolate a single query or mutation and work through its edge cases. Others walk a feature the way a user would, across several operations:
describe('Publishing an article', () => {
let app: INestApplication;
let graphql: GraphQLClient;
beforeAll(async () => {
const module = await buildAppModuleBuilder().compile();
app = module.createNestApplication();
await app.init();
registerDataSource(app.get(DataSource));
graphql = new GraphQLClient(app);
});
afterAll(() => app.close());
it('publishes a draft exactly once', async () => {
const draft = await graphql.createDraft({ title: faker.lorem.sentence() });
const first = await graphql.publishArticle(draft.id);
const second = await graphql.publishArticle(draft.id);
expect(first.errors).toBeUndefined();
expect(first.data.publishArticle.publishedAt).not.toBeNull();
expect(second.errors?.[0].extensions).toEqual(
expect.objectContaining({ code: 'ARTICLE_ALREADY_PUBLISHED' }),
);
});
});
Check errors before you read data, every time. GraphQL returns a 200 when a resolver throws, and a test that only reads data fails with a confusing undefined instead of the actual error.
Consumers are part of the surface too. Their specs boot the consumer’s module against the real RabbitMQ container, publish through a real publisher, and poll the database until the effect lands:
await publisher.publish({ articleId });
const article = await waitFor(
() => articles.findOneByOrFail({ id: articleId }),
(found) => found.indexedAt !== null,
{ timeoutMs: 10_000 },
);
expect(article.indexedAt).toBeInstanceOf(Date);
The harness: compose and a bash script
None of this is worth much if running it takes a wiki page. Each suite has one entrypoint: a compose file and a bash script that stands up a fresh database, migrates it and runs Jest.
# docker-compose.integration.yml
services:
postgres:
image: postgres:16
environment:
POSTGRES_USER: test
POSTGRES_PASSWORD: test
POSTGRES_DB: app_test
ports:
- '5434:5432'
tmpfs:
- /var/lib/postgresql/data
healthcheck:
test: ['CMD-SHELL', 'pg_isready -U test -d app_test']
interval: 2s
timeout: 5s
retries: 10
wiremock:
image: wiremock/wiremock:latest
ports:
- '8080:8080'
#!/bin/bash
set -e
COMPOSE_FILE="docker-compose.integration.yml"
if [ "$1" = "down" ]; then
docker compose -f "$COMPOSE_FILE" down -v
exit 0
fi
SETUP_ONLY=false
if [ "$1" = "--setup" ]; then
SETUP_ONLY=true
shift
fi
set -a
source .test.env
DB_PORT=5434
set +a
# A fresh database, every run.
docker compose -f "$COMPOSE_FILE" down -v 2>/dev/null
docker compose -f "$COMPOSE_FILE" up -d --wait
yarn build
# Under ARM emulation the healthcheck can pass before Postgres accepts
# connections, so migrations get a few attempts.
for i in 1 2 3 4 5; do
npx typeorm migration:run -d dist/data/data-source.js && break
[ $i -lt 5 ] || { echo "Migrations failed after 5 attempts."; exit 1; }
sleep 3
done
[ "$SETUP_ONLY" = "true" ] && exit 0
npx jest --config ./test/jest-integration.json --runInBand "$@"
A few of those decisions are load-bearing. Postgres runs on tmpfs, so the data directory lives in memory and a fresh database costs seconds. The E2E and integration suites get separate compose files on separate ports, 5433 and 5434, so both can run at once, locally or as parallel CI jobs. --runInBand because every suite in a run shares one database. And --setup stops after migrations, so CI can run setup and tests as separate steps with separate timings.
One trap has bitten more than once. Migrations run from the compiled dist/, and an incremental build never deletes a file whose source has gone. Switch branches and a migration that only exists on the other branch keeps running against your test database, cheerfully deleting seed rows your tests depend on. The symptom is an EntityNotFoundError against a lookup table that the migration log swears is fine. Delete dist/ before the first run on a new branch.
The subgraph goes a step further and runs Jest itself inside a container on the compose network, so tests address database and mockserver by service name. Cold starts are slower, but a run on my laptop, in CI and inside a dev container behaves identically.
Handing the harness to an agent
Everything above was worth doing before coding agents. It’s worth a lot more now. I wrote in the 2026 stack post that the unlock wasn’t the model, it was scaffolding that lets the model verify its own work, and this is that scaffolding. Once a test runs from one command, an agent can write a change, run the relevant slice, read the failure and try again without me in the loop.
What it needs next is to be told how. Both codebases carry Claude Code skills: markdown files with a description that decides when they load and a body that says what to do. The most valuable ones are the least glamorous. This is the gist of the one for running integration tests:
---
name: run-integration-tests
description: Use whenever you need to run, scope, or debug integration tests
(test/integration/**/*.spec.ts). They cannot run with a bare `jest` on the
host, only through ./scripts/integration.sh, which recreates the database,
runs migrations and then runs Jest.
---
# Run integration tests
Always use the script, and always scope it:
bash scripts/integration.sh article.repository
bash scripts/integration.sh -t "excludes archived articles by default"
Cold runs take minutes. Run in the background and wait for completion rather
than polling.
## Why not jest directly
Run without the script and you get `ENOTFOUND database`, hook timeouts, or
`relation "articles" does not exist`. Those are wiring failures, not code bugs.
Don't debug them, and don't export DB_HOST or edit the Jest setup to force a
host run. Use the script.
## On failure
An assertion or constraint error is a real bug: fix the code or the test. An
image build failure, a port already allocated or a failed migration is the
environment: report it rather than bypassing the harness.
That skill exists because of one specific, repeated failure. Without it, the agent reaches for jest directly, gets ENOTFOUND database, decides that’s a bug, and can lose an hour “fixing” environment wiring that was never broken. The skill names the symptoms so the agent recognises them, explains why they aren’t bugs, and closes off the escape hatches it would otherwise try.
The others follow the same pattern: state the rule, show the code, explain the failure it prevents. One enforces the service folder and the options pattern. One covers loaders: a single query per batch, key order, cache off. One covers list queries, and ends with a mandatory index audit of every column in a new WHERE or ORDER BY, with instructions to stop and ask when a filter can’t be helped by an index, like a leading-wildcard ILIKE. The migration skill is covered below.
The “why” is the part that matters. A rule on its own gets followed literally, including in the case it was never written for. A rule with its reason gives the agent enough to notice when it doesn’t apply, and to say so.
A few more pieces round it out. Each test directory has its own AGENTS.md with mocking rules, because they differ by layer: never mock the database in integration tests, always mock what you don’t own in E2E, mock freely in unit tests. The subgraph ships a dev container with Docker-in-Docker, so the compose-based harness runs unchanged inside it, and Claude’s permissions are relaxed only in there: edits auto-approved, the test scripts allow-listed, anything destructive or external still prompting. That’s the “agents get their own machine” direction from the stack post, and I think it’s where this is heading.
On the review side, the CRM has a GitHub Action that runs Claude Code when someone mentions @claude on a pull request, with a prompt carrying the project’s own checklist: every resolver guarded, UUID primary keys on new tables, migrations idempotent and non-breaking, truncation lists kept in step with new tables, field resolvers going through loaders. It’s a second reviewer that has actually read the conventions. It doesn’t make human review optional, but the obvious misses are mostly gone by the time a human looks.
Commands with nest-commander
Every project gets a CLI entrypoint beside main.ts, built on nest-commander so commands are ordinary providers with the same injected services as everything else:
// cli.ts
import { CommandFactory } from 'nest-commander';
import { CLIModule } from './modules/cli/cli.module';
async function bootstrap() {
await CommandFactory.run(CLIModule, ['warn', 'error']);
// An open database pool or broker connection would otherwise keep the
// process alive after the command has finished.
process.exit(0);
}
bootstrap();
@Command({
name: 'articles:reindex',
description: 'Push every published article back into the search index.',
})
export class ReindexArticlesCommand extends CommandRunner {
constructor(
private readonly articles: ArticleRepository,
private readonly search: ArticleSearchService,
) {
super();
}
public async run(_args: string[], options?: { apply?: boolean }): Promise<void> {
const ids = await this.articles.findPublishedIds();
if (!options?.apply) {
console.log(`Dry run: ${ids.length} article(s) would be reindexed. Re-run with --apply.`);
return;
}
await this.search.index(ids);
}
@Option({ flags: '--apply', description: 'Write changes instead of printing the plan.' })
public parseApply(): boolean {
return true;
}
}
That’s the shape of every command that writes: a dry run by default that prints the plan, and --apply to act on it. The ones that write somewhere expensive go further. The CRM’s identity commands print a banner naming the environment, with the organisation’s name fetched live from the provider rather than echoed from config, and then require the environment’s name to be typed back before anything happens. Not y/n, because the risk is muscle memory: someone who has run it against staging ten times pressing y on the eleventh without reading. The prompt is skipped when stdin isn’t a TTY, so CI isn’t left hanging, and an unattended run’s protection is the banner in its logs.
Commands are also how I prefer to run scheduled work. @nestjs/schedule is convenient, but an in-process cron runs on every instance of the API, so a job that must run once needs a lock or it runs once per instance. A command, or a token-guarded endpoint, triggered by the platform’s scheduler (a Kubernetes CronJob, a scheduled GitHub Action) runs once, with its own logs and its own exit code.
Consumers: same codebase, another entrypoint
Queue consumers live in the same repository and the same build as the API, behind a different entrypoint. That’s the whole trick. A consumer that reindexes an article when it’s published needs the same repositories, validation and business rules the API used to publish it. In a separate repository those get duplicated, or extracted into a shared package that every change now has to be versioned through. In the same codebase, the consumer imports the module.
The subgraph gives each consumer its own entrypoint and its own deployment, with a shared bootstrap doing the work:
// consumers/bootstrap.ts
import '../tracer';
import { DynamicModule, FactoryProvider, Module } from '@nestjs/common';
import { NestFactory } from '@nestjs/core';
import { AppModule } from '../modules/app/app.module';
import { MessengerModule } from '../modules/messenger/messenger.module';
import { Logger } from '../modules/observability/components/logger/logger.component';
@Module({})
class ConsumerModule {
static register(provider: FactoryProvider): DynamicModule {
return {
module: ConsumerModule,
imports: [AppModule, MessengerModule],
providers: [provider],
};
}
}
export async function bootstrap({ provider }: { readonly provider: FactoryProvider }) {
// An application context: the full container, no HTTP listener.
const app = await NestFactory.createApplicationContext(ConsumerModule.register(provider), {
bufferLogs: true,
});
const logger = app.get(Logger);
app.useLogger(logger);
app.enableShutdownHooks();
try {
await app.get(provider.provide).start();
} catch (error) {
logger.error('Failed starting consumer.', error);
throw error;
}
}
Each consumer then needs a provider, a three-line main.ts, and a Nest CLI config naming that file as its entry:
// consumers/article-published/main.ts
import { bootstrap } from '../bootstrap';
import { ArticlePublishedConsumerProvider } from './article-published.consumer.provider';
bootstrap({ provider: ArticlePublishedConsumerProvider });
{
"collection": "@nestjs/schematics",
"sourceRoot": "src",
"entryFile": "consumers/article-published/main"
}
"start:dev:consumer:article-published": "nest start --config src/consumers/article-published/config.json --watch",
"start:prod:consumer:article-published": "node dist/consumers/article-published/main.js"
The consumer class stays small, because the real work is a service call it shares with the API:
@Injectable()
export class ArticlePublishedConsumer {
constructor(
private readonly messenger: MessengerService,
private readonly search: ArticleSearchService,
) {}
public async start(): Promise<void> {
await this.messenger.addConsumer<ArticlePublished>(
{
exchange: { name: 'articles', routingKey: 'article.published' },
queue: { name: 'search.article-published' },
},
async ({ articleId }) => {
await this.search.index([articleId]);
},
);
}
}
The CRM takes the simpler route: one worker entrypoint that registers every consumer in onModuleInit, deployed as one extra image built from the same source. A deployment per consumer buys independent scaling and failure isolation; a single worker buys less infrastructure. For a handful of queues the worker is the right call. For a dozen, some hot and some idle, separate deployments are.
Either way, the topology is asserted on startup: a durable topic exchange, a dead-letter exchange and queue beside it, and the consumer’s queue bound with x-dead-letter-exchange set. A handler that throws never has its message requeued in place, so one poison message can’t hot-loop. The CRM sends it straight to the dead-letter queue; the subgraph’s consumers retry with exponential backoff first. Either way, the dead-letter queue is somewhere to inspect and replay from, and its depth is a monitor.
The honest caveat: in both codebases, consumers still import the whole AppModule, GraphQL module included. Nothing listens, so it’s harmless, but every consumer builds a schema it’ll never serve on every boot. In the subgraph that should be a small change, because the GraphQL layer is already its own module. In the CRM, with the schema bolted onto the entities, it’s a much longer job. That’s the argument from the structure section, coming due.
Migrations
TypeORM migrations, and never synchronize. The rules for writing them have hardened over the years into a skill:
- Raw SQL through
queryRunner.query(), with nothing imported from the application. Entities and enums get renamed, and a migration that imports them breaks on a fresh database years later. - Idempotent, using
IF NOT EXISTS,IF EXISTSandON CONFLICT DO NOTHING, so a partially applied migration can simply run again. - The timestamp comes from the real clock, unless an older migration carries a future-dated one, in which case it’s one past the highest. TypeORM orders by that number, and a migration that sorts before its predecessors breaks on a fresh database.
- Non-breaking, always. A nullable column, a new table or a new index is fine. Dropping, renaming, narrowing a type or tightening nullability is a two-release job: ship the code that stops using the thing, then ship the migration that removes it.
That last rule follows from how they run. Migrations run in CI as their own job, minutes ahead of the deploy, so the currently running version of the app is the first thing to meet the new schema. Drop a column the live code still reads, and production errors until the new release rolls out.
import { MigrationInterface, QueryRunner } from 'typeorm';
/**
* Lets editors schedule an article for a future publish time. Nullable, so
* existing rows and the currently deployed code are unaffected.
*/
export class ArticlePublishAt1789975178299 implements MigrationInterface {
public async up(runner: QueryRunner): Promise<void> {
await runner.query(`
ALTER TABLE articles
ADD COLUMN IF NOT EXISTS publish_at TIMESTAMPTZ NULL;
CREATE INDEX IF NOT EXISTS articles_publish_at_idx
ON articles (publish_at)
WHERE publish_at IS NOT NULL;
`);
}
public async down(): Promise<void> {}
}
The empty down() is deliberate in the CRM: it never rolls a migration back, so reversal SQL would be effort spent on a path nobody takes. The subgraph writes idempotent down() methods anyway. Forward-only is the more honest of the two. If a migration is wrong, the fix is another migration.
On the CI side, migrations are a dedicated GitHub Actions job, and the deploy can’t start until it’s green:
jobs:
# build and test jobs elided
database-migrations:
runs-on: ubuntu-latest
name: TypeORM Migrations
environment: production
needs: [build-server]
env:
DB_HOST: ${{ secrets.DB_HOST }}
DB_PORT: ${{ secrets.DB_PORT }}
DB_NAME: ${{ secrets.DB_NAME }}
DB_USERNAME: ${{ secrets.DB_USERNAME }}
DB_PASSWORD: ${{ secrets.DB_PASSWORD }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: '24'
- run: yarn install --immutable
- run: yarn build && npx typeorm migration:run -d dist/data/data-source.js
deploy:
runs-on: ubuntu-latest
environment: production
needs:
- database-migrations
- workos-authorization
steps:
- run: ./scripts/deploy.sh production
environment scopes the database credentials to the jobs that need them, and lets a protection rule gate a production release. deploy waiting on database-migrations is the whole point: the schema always arrives before the code that needs it. The pull request workflow runs the same job against staging once a PR’s tests pass, which is one more reason every migration has to be safe for code that doesn’t know about it yet. The second job deploy waits on is next.
Authentication with WorkOS
Authentication is the part of these projects that has changed most recently. The CRM moved staff sign-in to WorkOS AuthKit this year, and the application no longer stores a password, generates a TOTP secret or issues its own tokens. Sign-in, MFA enrolment, password resets and session refresh all belong to WorkOS. I’d been writing some version of that flow on every client project for years, and AuthKit has saved me a lot of it.
The token, the principal and the local row
The API verifies AuthKit’s RS256 access tokens with passport-jwt against the WorkOS JWKS, fetched and cached by jwks-rsa so a key rotation needs no deploy. The issuer and JWKS URL are both derived from the client id, which is what makes a token minted by the staging environment fail signature verification in production instead of being trusted for whatever role it claims.
What lands on the request is a principal: the token’s claims, plus the local user row.
export type Principal = {
readonly workosUserId: string;
readonly organizationId: string;
readonly role: Role;
readonly permissions: readonly Permission[];
/** The local row every foreign key in the system points at. */
readonly user: User;
};
The local users table doesn’t go away, because every foreign key in the system points at it: who wrote the article, who moderated the comment, who changed the setting. WorkOS owns identity, the local row is what the data hangs off, and the two are linked by a WorkOS user id on the row and the row’s id stored as the WorkOS account’s external id.
One service builds the principal, and it’s the only place that decides what happens when the token and the database disagree. It refuses a token with no organisation, a token from a different organisation, a role the application doesn’t declare, and an inactive user. On a miss it reads through, claiming an existing row by email or creating one, so somebody invited from the WorkOS dashboard can sign in without anyone running a command. @CurrentUser() still returns the local row, which is why roughly ninety resolvers didn’t change when the auth underneath them did.
Even a single-tenant app needs one WorkOS organisation with everyone as a member. WorkOS only puts role and permissions claims into an organisation-scoped session, so without a membership there’s nothing to authorise with.
Users, both directions
The client wanted to keep managing users inside the CRM rather than learn a second dashboard, so user management runs both ways.
Outbound, the create and update mutations write the local row first, then forward the identity half to WorkOS: an invitation, a name, a role change on the membership, an activation or deactivation. The order is deliberate, because the two failure modes aren’t symmetric. Row-then-WorkOS fails to a user who exists locally but can’t sign in yet, who shows up on the users screen and is fixed by saving the form again. WorkOS-then-row fails to an account that can authenticate against an application that has never heard of it. Every step is idempotent so that saving again really does finish the job.
Inbound, there are three writers. A webhook is the fast path, verified by HMAC over the raw request body (Nest’s rawBody: true), because re-serialising the parsed body doesn’t reliably reproduce the bytes that were signed. A workos:sync command is the backstop, following the Events API from a stored cursor, with a --full mode that sweeps memberships directly. That mode exists because an event stream with a retention window is a change feed, not a backlog, and an incremental first run against an empty table looks like a success while mirroring nothing. The webhook and the sync both call the same applyEvent, so a push and a pull of one event can’t diverge. And the principal service’s read-through corrects drift for whoever is making the current request.
Deactivation never deletes. The row is referenced by years of history, so a deleted WorkOS user becomes an inactive local one with its WorkOS id cleared.
One smaller thing from this work I now do everywhere. The self-service profile mutation takes no user id and its input type has no field that grants anything. There’s no subject argument to aim it at someone else’s account, and no role or active field to escalate with. The input type is the authorisation, and an E2E test asserts the privileged fields are rejected so nobody adds one back without noticing.
Roles and permissions as code
This is the part I’d push hardest on any team adopting WorkOS. The application owns its roles, its permissions and the grants between them. WorkOS just stores them. The manifest is a file in the repository:
export const PERMISSIONS = [
'articles:publish',
'articles:delete',
'comments:moderate',
'authors:manage',
'users:manage',
'settings:manage',
] as const;
export type Permission = (typeof PERMISSIONS)[number];
export const ROLE_PERMISSIONS: Readonly<Record<Role, readonly Permission[]>> = {
writer: [],
editor: ['articles:publish', 'comments:moderate'],
admin: [...PERMISSIONS],
};
A guard reads the permissions claim straight off the token:
@Mutation(() => Article)
@UseGuards(GraphAuthGuard, PermissionGuard.of('articles:publish'))
public publishArticle(@Args('id', { type: () => ID }) id: string): Promise<Article> {
return this.articleService.publish(id);
}
A workos:push-authorization command makes the WorkOS environment match the manifest. It runs in CI as its own job next to the migrations, and deploy waits on both:
workos-authorization:
runs-on: ubuntu-latest
name: WorkOS Roles & Permissions
environment: production
needs: [build-server]
env:
WORKOS_API_KEY: ${{ secrets.WORKOS_API_KEY }}
WORKOS_CLIENT_ID: ${{ vars.WORKOS_CLIENT_ID }}
WORKOS_ORGANIZATION_ID: ${{ vars.WORKOS_ORGANIZATION_ID }}
WORKOS_ENVIRONMENT: ${{ vars.WORKOS_ENVIRONMENT }}
# Plus the database variables: the CLI boots the full application.
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: '24'
- run: yarn install --immutable
- run: yarn cli workos:push-authorization --apply
A release that introduces a permission has to create it before the code that checks it serves traffic. A guard on a slug WorkOS has never heard of returns 403 for everybody, which from the outside looks like a broken feature rather than a missing object. With the job in the pipeline, a change that needs a new permission declares it, grants it and checks it in one commit, and reviewing the manifest is reviewing the policy. The client id, organisation id and environment label are variables rather than secrets on purpose: GitHub redacts secrets from logs, and those three are what the command’s banner prints to show which environment an unattended apply is pointed at.
The sync’s boundaries took some thought:
- It creates and updates roles and permissions, and converges each declared role’s grants, which means it revokes. An add-only sync can never take a permission away, and a matrix that can only grow isn’t a source of truth.
- It only revokes slugs the manifest declares. Anything else on a role is copied through untouched, and a role the manifest doesn’t declare is never touched at all.
- It never deletes. WorkOS won’t rename a role slug and its API can’t delete one, so a new role is a one-way door anyway.
- It doesn’t run on pull requests. Unlike a migration it converges, so two open PRs that disagree about a grant would flip a shared staging environment back and forth between them.
The last guard rail came from a limit that’s easy to miss. WorkOS puts the signed-in role’s entire permission list into the access token and won’t issue one over 3 KB, so the manifest has a size budget, and the admin role, holding every permission, is what spends it. Blow the budget and sign-in fails at the code exchange, after WorkOS has already recorded a successful authentication, for every holder of that role at once. The push command now weighs each role’s permission list before writing anything and refuses to apply if one won’t fit. A red deploy is a far better outcome than a locked-out admin team.
Monitoring
Datadog, in all of these. Most of the work is making sure every signal carries enough context to be useful.
The tracer is initialised in its own file and imported on the first line of every entrypoint, before anything it needs to instrument has loaded:
// tracer.ts
import tracer from 'dd-trace';
tracer.init({ logInjection: true, runtimeMetrics: true });
tracer.use('graphql', {
depth: 2,
hooks: {
execute: (span, _args, result) => {
const statusCode = result?.errors?.[0]?.extensions?.statusCode;
// A validation failure or a missing record is the client's problem, not
// an outage. Only 5xx-class errors count against the error rate.
if (span && typeof statusCode === 'number' && statusCode < 500) {
span.setTag('error', false);
}
},
},
});
export default tracer;
That hook fixes a quirk that would otherwise make the error rate useless. A resolver that rejects bad input produces a GraphQL error, the integration marks the span as failed, and every bad request from every client reads as a server fault. A global exception filter maps domain exceptions onto GraphQL errors with a stable code and a statusCode in the extensions, logging 5xx as errors and 4xx as warnings, and the hook un-flags anything below 500. The error rate goes back to meaning “we broke something”.
Logs go through a custom Logger that implements Nest’s LoggerService over winston, writes JSON in deployed environments, and attaches the context of the request that produced each line. A middleware opens an AsyncLocalStorage store per request, and the logger reads it on every call:
@Global()
@Module({
providers: [LoggerProvider, Metrics, AsyncLocalStorageProvider],
exports: [Logger, Metrics],
})
export class ObservabilityModule implements NestModule {
constructor(private readonly store: AsyncLocalStorage<ObservabilityStore>) {}
public configure(consumer: MiddlewareConsumer) {
consumer
.apply((req: Request, _res: Response, next: NextFunction) => {
const context: HttpObservabilityStore = {
request: {
id: firstHeader(req, 'x-request-id', 'cf-ray') ?? randomBytes(8).toString('hex'),
endpoint: req.baseUrl,
client: {
name: firstHeader(req, 'apollographql-client-name') ?? 'unknown',
version: firstHeader(req, 'apollographql-client-version'),
},
auth: extractAuth(req), // decoded, not verified: for logs only
gql: extractGqlOperation(req), // operation name parsed from the body
},
};
this.store.run(context, () => next());
})
.forRoutes('*');
}
}
The request id comes from x-request-id or Cloudflare’s cf-ray before falling back to random, so a log line can be followed back through the edge. The client name and version come from Apollo’s client headers, which answers “which app version is sending this?” without a deploy. The auth block is the token decoded but not verified, purely for the logs; verification belongs to the guard. The operation name is parsed out of the GraphQL body, because a hundred log lines that all say POST /graphql tell you nothing.
Metrics go through a small Metrics provider wrapping dd-trace’s DogStatsD client, and the most useful ones come from ObservableDataLoader:
export class ObservableDataLoader<K, V> extends DataLoader<K, V> {
constructor(batchLoadFn: DataLoader.BatchLoadFn<K, V>, options: Options<K, V>) {
super(
(keys) => {
const tags = {
'dataloader.name': options.name ?? this.constructor.name,
'dataloader.bucket': getBatchSizeBucketName(keys.length),
};
options.metricsAgent.increment('dataloader.batch-load', 1, tags);
options.metricsAgent.increment('dataloader.batch-load.key-count', keys.length, tags);
return tracer.trace('dataloader.batch-load', () => batchLoadFn(keys));
},
{ cache: false, ...options },
);
}
}
export function getBatchSizeBucketName(size: number) {
if (size === 1) return 'single';
if (size <= 5) return 'xs';
if (size <= 10) return 'sm';
if (size <= 25) return 'md';
if (size <= 50) return 'lg';
if (size <= 100) return 'xl';
return '2xl';
}
The bucket keeps tag cardinality fixed, where tagging the raw batch size would mint a new tag value for every size. What it buys is a dashboard of how well each loader actually batches. A loader whose traffic suddenly lands in single has lost its batching, and the usual cause is somebody awaiting inside a loop.
Health is split in two. Liveness returns ok as long as the process can respond at all. Readiness runs SELECT 1 against the database and checks the broker connection’s healthy flag, and returns a 503 if either is down, so an instance that has lost its broker drops out of rotation instead of accepting work it can’t finish.
Monitors are Terraform, living in the same repository as the code they watch, so a new consumer and its alerts can arrive in the same change. Every queue gets two as standard: no consumers attached, and dead-letter depth over a threshold.
resource "datadog_monitor" "article_published_no_consumers" {
count = var.environment == "prod" ? 1 : 0
name = "[Articles] No consumers on search.article-published (${var.environment})"
type = "query alert"
query = "avg(last_1h):sum:rabbitmq.queue.consumers{env:${var.environment},rabbitmq_queue:search.article-published} <= 0"
monitor_thresholds {
critical = 0
critical_recovery = 0.1
}
message = "Nothing is consuming search.article-published on ${var.environment}. ${local.notify}"
tags = ["service:articles", "terraform:true", "env:${var.environment}"]
}
Rate-based monitors get one more trick, straight from Datadog’s own guide: multiply the failure ratio by is_greater(total, 20), so the alert can only fire once there’s a minimum sample. A 50% failure rate across two requests at 3am isn’t an incident.
What I’d still change
Both codebases pass most of the rules in this post. Neither passes all of them.
The CRM’s entities still double as GraphQL types, and every new feature makes the split more expensive. Both codebases boot the entire application, GraphQL included, inside consumers and commands that will never serve a request. The CRM still runs a couple of scheduled jobs in-process on every API instance. And its resolvers still authorise by role rather than by permission, even though the permissions now exist in WorkOS; moving forty call sites is its own reviewed change, and it hasn’t happened yet.
None of those are hard. They’re just never urgent, which in a codebase amounts to the same thing.