The system works. Until it stops
Imagine an online store where customers can browse products, add them to the cart, and proceed to checkout. The server responds, the page loads, and the basic infrastructure metrics show no serious problem. At first glance, everything seems to work correctly.
Meanwhile, some customers cannot complete their orders. For some, payment takes several seconds; for others, an error appears. The technical team receives reports, but does not yet know whether the cause is the payment gateway, the database, the latest application update, or perhaps a communication problem between services.
This is a situation in which ordinary monitoring may prove insufficient. It can show that the number of errors has increased or response times have lengthened, but it does not always provide the information needed to quickly find the cause.
This is exactly the gap that observability addresses, or system observability. Its goal is not merely to state that the application is not working properly. It is about the ability to analyze its behavior, reconstruct the sequence of events, and find the source of the problem, even when a specific failure scenario was not anticipated in advance.
Monitoring and observability - similar goals, different capabilities
Monitoring and observability are closely related, but they do not mean the same thing.
Monitoring consists of systematically collecting and analyzing data about the state of the application and infrastructure. It allows you to track defined parameters, detect deviations from accepted norms, and trigger alerts when a situation requiring action occurs.
Examples of questions that monitoring answers include:
- Is the server available?
- What is the average API response time?
- Is the number of HTTP 500 errors exceeding the set threshold?
- What is the memory and CPU usage?
- Is the task queue growing faster than the system can handle it?
Observability goes further. It allows you to analyze data from different parts of the system and combine it into context that helps explain why a particular behavior occurred.
This can be captured in three questions:
- Monitoring: Is something wrong?
- Diagnostics: Where did the problem appear?
- Observability: What happened, why, and what was the impact on system operation?
Observability does not replace monitoring. It is an approach that uses monitoring and properly prepared telemetry data to enable deeper investigation of application behavior. OpenTelemetry describes observability as the ability to ask a system questions about its behavior based on signals such as logs, metrics, and distributed traces. <Cite ref="turn154932search0"/>
The three pillars of observability: logs, metrics, and tracing
The foundation of observability consists of three types of telemetry data: logs, metrics, and traces. Each shows a different aspect of how the application works. Only their correlation provides a broader picture of the situation.
1. Logs - what happened in the application?
Logs are structured records of events generated by the application, operating system, servers, databases, and other infrastructure components.
They can record, for example:
- the start and end of a process,
- a user login attempt,
- a request sent to an external API,
- a data validation error,
- a failed transaction,
- an exception raised by the application,
- a change in order status.
A log can contain a timestamp, severity level, service name, message, request ID, and additional attributes describing the event.
It is worth distinguishing text logs from structured logs. A text entry may look like this:Error while processing the order
Such a message informs you that a problem occurred, but says little about its context. A structured log, however, may contain separate fields such as order ID, operation name, error code, execution time, and trace ID. This allows the data to be filtered, grouped, and analyzed automatically.
Best practice: logs should be designed with future analysis in mind, not just for storing messages. It is worth using consistent naming, severity levels, and correlation identifiers that make it possible to link events from different components.
At the same time, logging requires common sense. Recording every operation in full can generate huge amounts of data, increase storage costs, and make it harder to find relevant information. Equally important, logs should not reveal passwords, access tokens, payment card data, or other confidential information. Masking, access control, appropriate retention periods, and secure data processing practices should be used.
2. Metrics - how does the system behave over time?
Metrics are aggregated numerical data describing the state, performance, and behavior of a system over a given period of time.
Example metrics include:
- the number of requests handled per second,
- the percentage of requests ending in error,
- API response time,
- CPU and memory usage,
- the number of active sessions,
- the length of the task queue,
- the number of completed transactions,
- the wait time for a database connection.
Their biggest advantage is the ability to observe trends. A single error may be an incident, but a gradually increasing response time, a growing number of failed transactions, or a swelling task queue may indicate a developing problem.
In practice, metrics help answer not only whether the application is working, but also whether its performance meets user expectations and business requirements.
Response time percentiles, such as p95 and p99, are particularly useful. An average can hide a situation where most users get a quick response, but a small portion experiences very long delays. Percentiles show how long it takes to handle slower requests and help reveal problems invisible in average values.
Best practice: choose metrics that matter to the user and the business process. Information about CPU load alone will not tell you whether a customer can place an order. It is therefore worth monitoring indicators related to key functions as well, such as payment success rate, order fulfillment time, or the availability of the most important operations.
3. Tracing - where did the request go?
Tracing, and especially distributed tracing, allows you to follow the path of a single request through different components of the system.
In a modern application, a user action can involve many stages. Clicking the “Place order” button can trigger a request in the browser that goes to the API, then to the order service, the database, the warehouse system, and an external payment provider.
If the process is delayed, the response time of the entire API alone may not be enough. Tracing makes it possible to break this time down into individual operations and see which stage is responsible for the delay.
The basic element of a trace is a span, which is a record of a single operation. A span can contain the start and end time, the name of the operation, status, and metadata. Related spans form a trace showing the path of the entire request.
For example:
- The API receives the request and passes it on.
- The order service validates the data.
- The database saves the order.
- The warehouse service checks product availability.
- An external payment gateway processes the transaction.
- The application returns the result to the user.
If the whole process takes 8 seconds, tracing can show that 6.5 seconds were spent waiting for the external payment gateway’s response, while the other operations proceeded correctly. The team then gets a concrete starting point for further analysis instead of beginning diagnostics with random components.
Tracing is especially useful in microservice architectures, distributed systems, queue-based applications, and solutions integrating many external services. <Cite ref="turn154932search0"/>
Alerts - information must reach the right person
Telemetry data is useful only when it can be acted upon. That is why an alerting system is an important part of observability.
An alert is a notification about an event or condition that requires attention. It may be triggered when a metric threshold is exceeded, a specific error pattern is detected, or it is determined that a key application function is not working as expected.
However, not every increase in load should generate an alarm. If the system regularly handles heavy traffic during peak hours, notifying on every increase in request volume will create noise. Too many alerts lead to them being ignored and, as a result, a truly important incident may be missed.
It is therefore worth determining:
- which events require immediate reaction,
- which problems can be analyzed in standard operating mode,
- who is responsible for a given type of alert,
- what information a notification should contain,
- what actions should be taken after receiving it.
A good starting point is to define alerts based on user impact and reliability goals, rather than solely on infrastructure parameters. For example, an alert about an increasing share of failed payments may be more important to the business than a short-lived spike in CPU usage.
An alert should lead to action. If it is not clear who should handle it or what needs to be done, it is merely another message in the system.
A practical example: how does observability help find the cause of a failure?
Let’s assume that users of a B2B application report that report generation takes much longer than usual. Monitoring detects an increase in response time and triggers an alert.
The team begins the analysis:
- Metrics show that the problem mainly affects reports covering large data ranges. Other functions are working normally.
- Tracing shows that the biggest delay occurs during a database query.
- Logs contain details of the query, its operational parameters, and error information without exposing sensitive data.
- Data correlation makes it possible to link a specific trace with the relevant log entries and changes visible on metric charts.
- Change analysis shows that the problem appeared after deploying a new version of the report, which began executing an expensive query.
As a result, the team does not have to inspect the entire infrastructure blindly. It can focus on a specific operation, compare behavior before and after deployment, and then optimize the query or roll back the change.
Observability does not eliminate failures and does not guarantee that every cause will be found automatically. However, it helps narrow the search area, shorten diagnosis time, and base decisions on data rather than assumptions.
Data correlation - the greatest value appears together
Logs, metrics, and tracing are useful on their own, but their real value emerges when they can be linked together.
Imagine that a dashboard shows a sudden increase in response time. A metric indicates when and to what extent the problem appeared. A trace shows which operations made up the slow request. Logs allow you to check what events occurred at a specific stage.
To make this possible, the system should consistently pass request context between services. Trace and span identifiers can be used to link log entries with traces. It is also worth keeping consistent information about the service name, environment, application version, and other attributes describing the data source.
Without correlation, the team may have access to many dashboards, files, and tools, yet still waste time manually determining which events are related. OpenTelemetry identifies the correlation of logs, traces, and resource context as an important element of building useful telemetry. <Cite ref="turn154932search1"/>
OpenTelemetry - a common standard for telemetry data
Implementing observability does not have to mean dependence on a single tool vendor. One solution that supports interoperability is OpenTelemetry (OTel) - an open set of standards, APIs, libraries, and tools for instrumentation, generation, collection, and export of telemetry data.
OpenTelemetry enables an application to emit metrics, logs, and traces in a consistent model. The data can then be sent to a chosen observability backend responsible for storage, search, visualization, and analysis.
An important part of the ecosystem is the OpenTelemetry Collector. It can receive data from various sources, process it, enrich it with additional context, and export it to configured systems. As a result, the application does not have to be directly tied to every tool used for analysis.
This approach is especially useful when a company uses many technologies, is evolving its system architecture, or wants to keep the option to change tool vendors. However, the standard itself does not provide complete observability. Proper instrumentation, a thoughtful data collection strategy, appropriate dashboards, alerts, and response procedures are still needed. <Cite ref="turn154932search3"/>
When is it worth implementing observability?
Observability can be useful both in large distributed systems and in smaller applications where downtime or a hard-to-detect defect has significant business consequences.
This approach is particularly worth considering when:
- the application consists of many services or integrations,
- issues occur intermittently and are hard to reproduce,
- users report bugs that are not visible in standard tests,
- the time needed to diagnose incidents is too long,
- subsequent deployments cause hard-to-predict effects,
- the company is growing the system and needs data for performance planning,
- the application supports key sales, operational, or financial processes,
- the team needs a better understanding of the impact of external services on the operation of the entire solution.
This does not mean, however, that every website needs an advanced telemetry environment. In a small, simple service, basic logs, uptime monitoring, and a few key metrics may be enough. The scope of the solution should match the complexity of the application, traffic scale, reliability requirements, and the cost of potential downtime.
When can observability be overkill?
Deploying advanced tools without a clearly defined goal can bring more costs than benefits.
The most common mistakes include:
- Collecting everything without a plan. Too much data increases costs and makes it harder to find information relevant to diagnosis.
- Not asking what questions the system should answer. Dashboards may look impressive, but not help solve real problems.
- Alerting on every deviation. Too many notifications cause alert fatigue and increase the risk of missing an incident.
- No accountability for response. Even a well-detected problem can last a long time if no one knows who should handle it.
- No data protection. Telemetry may contain sensitive information, user identifiers, or operational data that require restricted access and appropriate retention.
- Ignoring the cost of instrumentation. Collecting detailed traces and logs at scale can affect application performance and generate significant storage and processing costs.
- Treating the tool as a ready-made solution. Simply installing a platform does not ensure proper instrumentation or an effective diagnosis process.
Observability therefore requires not only technology, but also organizational decisions: which data is needed, who analyzes it, how the team responds, and how the lessons from incidents translate into changes in the system.
How do you plan an observability implementation?
The safest approach is to develop observability in stages, starting with the processes and functions whose failure has the greatest impact on users and the business.
1. Define key business processes
Identify the most important operations, such as logging in, placing orders, payments, generating documents, or data synchronization. These should be the starting point for defining what proper application behavior means.
2. Establish reliability metrics
Choose metrics that reflect the user experience, for example the availability of key functions, response time, or the percentage of successfully completed operations. For important services, you can define SLI (Service Level Indicator), and SLO (Service Level Objective), meaning the target for that indicator.
3. Implement structured logs
Standardize the log format, severity levels, and basic attributes. Make sure correlation identifiers and a policy for deleting or masking confidential data are in place. Logs should be readable for the team and processable by tools.
4. Add tracing in critical paths
Start with processes that involve multiple services, databases, or external integrations. Track the request flow through the system and ensure context propagation between components.
5. Build dashboards around specific questions
Instead of creating one huge panel, prepare views that address the needs of different roles. The technical team may need information about errors and delays, while the product owner needs data on the effectiveness of key processes and the impact of failures on users.
6. Design alerts and response procedures
Define thresholds, priorities, responsible people, and instructions. An alert should contain context that helps start diagnosis quickly, not just notify about a threshold breach.
7. Test observability
Check whether the team can find the cause of a sample error using the available data. Controlled failure tests can be run in a test environment, or incident response exercises can be conducted. It is also worth verifying whether alerts trigger when they should.
8. Develop the solution based on incidents
After each significant issue, it is worth checking what information was available, what was missing, and how instrumentation, alerting, or procedures can be improved. Observability is not a one-time project, but a process of improving knowledge about how the system works.
Costs and security - two aspects that cannot be overlooked
Telemetry data comes at a price. Costs may arise from instrumentation, transfer, indexing, storage, retention, and data analysis. In high-traffic systems, collecting all traces or very detailed logs can be particularly expensive.
That is why it is worth using retention tailored to needs, data filtering, trace sampling, and different levels of detail depending on the environment. For example, a production system may collect full data for errors and selected critical operations, while sampling some successful requests to limit volume.
Data protection is equally important. Logs and traces can unintentionally contain personal data, session identifiers, fragments of queries, or information about infrastructure structure. Access to telemetry should be restricted, unnecessary data should be removed, sensitive information should be masked, and storage periods should be controlled. It is also worth treating observability systems as part of the production environment, which itself requires safeguards, backups, and permission control.
Observability as a management tool, not just for diagnostics
Although observability is most often associated with the work of developers, DevOps, and administrators, its value extends beyond IT.
Data on response times, errors, availability, and process effectiveness can help a company understand which elements of technology support the business and which hinder it. It makes it possible to identify recurring problems, assess the effects of changes, and plan development based on the system's actual behavior.
If, for example, the order system regularly slows down at specific hours, telemetry data can help determine whether query optimization is needed, whether the task processing method should be changed, or whether the infrastructure should be expanded. Instead of investing in additional resources based on intuition, the company can first identify the real bottleneck.
Observability also supports the analysis of deployment impacts. Comparing metrics, traces, and logs before and after a change makes it possible to detect regressions faster and assess whether the update delivered the expected result.
However, observability should not be equated with automatic decision-making. Data shows the behavior of the system, but interpreting it requires knowledge of the architecture, business processes, and the context of a specific incident.
Glossary
- Observability - the ability to understand a system's internal behavior based on the data it emits.
- Monitoring - the continuous tracking of selected parameters and the detection of specific states that require attention.
- Telemetry - data collected and transmitted from applications and infrastructure for the purpose of analyzing their operation.
- Logs - records of events occurring in an application or infrastructure.
- Metrics - numerical measurements of the state, performance, or behavior of a system over time.
- Tracing - tracking the path of operations through application components.
- Distributed tracing - tracking a single request in a system made up of many services or processes.
- Span - a record of a single operation that is part of a trace.
- Trace - a set of related spans showing the course of an operation.
- Alert - a notification about a detected state or event requiring a response.
- SLI - an indicator measuring a specific aspect of a service's operation.
- SLO - a defined target for a selected reliability indicator.
- Sampling - a technique for limiting the amount of telemetry data collected by selecting a representative portion of events.
- OpenTelemetry - an open set of standards and tools supporting the instrumentation and export of telemetry data.
Summary
An application failure does not always start with an unavailable server or an error message. Sometimes the system is formally working, but a key function becomes too slow, part of the transactions do not complete successfully, or an integration fails only under specific conditions.
Monitoring helps detect abnormalities. Observability makes it possible to understand what led to them and what impact they had on the application's operation. Logs, metrics, tracing, and well-designed alerts together form the foundation for more efficient diagnostics, informed development, and reduced operational risk.
The point is not to collect as much data as possible or create the most elaborate dashboards. The point is that when a problem occurs, we should not ask only: "Is the system working?", but be able to determine: "What exactly happened in the system, why, and what should we do next?"
A mature application is not only one that works. It is also one whose behavior can be understood, diagnosed, and improved.



