Skip to main content

2 posts tagged with "nestjs"

View All Tags

A Trace Id in Every Log Line: How Cross-Service Debugging Went From a Four-Engineer Session to One Grafana Query

· 14 min read
Ivan Baha
Software Team Lead & Architect

In 2022, I added a trace id to the two shared libraries every service my team owned already used – the logger and the HTTP client – so that one query in Grafana would return every log line of one request, across every service it touched. The platform is an enterprise system built as microservices from day one; it runs more than 200 of them today across several teams. My team owns 80+ of those. The trace id now covers more than 100 – more than we own – because other teams adopted it.

Before that change, a cross-service bug on a test environment was a meeting: a QA engineer, a business analyst, the team lead, usually a developer or two, sometimes DevOps – several hours of reading logs by timestamp and reconstructing the call chain from memory of the code. After that, the same investigation became a text field on a Grafana board, and the person typing into it was most often the analyst, not an engineer. Minutes rather than hours, and in most cases no engineer at all.

The cost is stated up front because it matters. There are no span IDs. Which service called which is inferred from a forwarded User-Agent and timestamps rather than being known. That was the trade.

This is a field report: what debugging looked like, what changed, and two cases as they happened. The mechanism is documented as a reference architecture, RA-004, and reconstructed as runnable code in team-workspace; this article explains it only as far as the story needs. If you read one section, read the two cases.

In-Process Authorization with Guards: The Default That's Enough Until It Isn't

· 6 min read
Vladyslava Prykhodko
Engineering Technical Lead & Architect

Once you have more than a handful of services, "can this caller do this thing to this resource?" stops being a one-liner. The answer usually depends on attributes — who the subject is, what they own, which tenant they belong to, what action they're attempting, sometimes the time of day. That's attribute-based access control, ABAC: the decision is a function of subject, resource, action, and context, rather than a flat list of roles.

The interesting question isn't whether to do ABAC. It's where the decision gets computed. There's a whole spectrum. This article is the cheap, simple end of it — and most teams should start, and often stay, here. Part 2 measures what the externalized alternative actually costs once you genuinely need it.

Everything below is illustrative and generalized — a reference model, not data or code from any specific system. Substitute your own.