When a gRPC call fails, the client gets back a status code. That code is the first clue to what went wrong — was the service down? Did the request time out? Is the method not implemented? The trouble is, there are 17 different gRPC status codes, and their meanings are tightly bound to when they should be used, who generates them (client library, server, proxy), and whether the client should retry. Understanding them is essential to error handling and monitoring in gRPC, and mistakes here can cause cascading failures across your microservices.
This guide explains all 17 status codes, shows you how to inspect them in code, and clarifies which ones are safe to retry and when the fault lies with the client versus the server.
All 17 gRPC status codes (reference table)
| Code | Name | Meaning | Retryable | Who Generates |
|---|---|---|---|---|
| 0 | OK | Success; no error. | No (success) | Library/server |
| 1 | CANCELLED | The operation was explicitly cancelled. | No | Client or middleware |
| 2 | UNKNOWN | An unspecified error occurred; often an unhandled server exception. | No | Server |
| 3 | INVALID_ARGUMENT | The client sent an invalid argument (malformed request, out-of-range value). | No | Client/server validation |
| 4 | DEADLINE_EXCEEDED | The deadline specified by the client was exceeded. | Yes (with caution) | Client library/server |
| 5 | NOT_FOUND | The requested resource does not exist. | No | Server |
| 6 | ALREADY_EXISTS | The operation tried to create something that already exists. | No | Server |
| 7 | PERMISSION_DENIED | The caller lacks the permission to execute the operation. | No | Server |
| 8 | RESOURCE_EXHAUSTED | The server ran out of a critical resource (memory, file descriptors, quota). | Yes | Server |
| 9 | FAILED_PRECONDITION | The operation was rejected because the system is not in the required state. | Yes | Server |
| 10 | ABORTED | The operation was aborted, typically due to a concurrency issue (conflict). | Yes | Server |
| 11 | OUT_OF_RANGE | The client tried to access data outside a valid range. | No | Server |
| 12 | UNIMPLEMENTED | The RPC method is not implemented on the server. | No | Server |
| 13 | INTERNAL | An unexpected internal server error; something went wrong inside the service. | No | Server |
| 14 | UNAVAILABLE | The service is temporarily unavailable, often due to shutdown, load shedding, or a transient failure. | Yes | Server/load balancer |
| 15 | DATA_LOSS | Unrecoverable data loss or corruption has occurred. | No | Server |
| 16 | UNAUTHENTICATED | The request lacks valid authentication credentials. | No | Authentication system |
The table tells you what each code means, but the story gets richer when you understand the context.
The retry family: UNAVAILABLE, DEADLINE_EXCEEDED, ABORTED, FAILED_PRECONDITION
These four codes are the most common causes of confusion — they all suggest "maybe try again." But the semantics differ, and retrying incorrectly can cause harm.
UNAVAILABLE (14)
The service is temporarily unable to handle requests. This includes:
- The server is shutting down and rejecting new work.
- A load balancer or service mesh detected the server as unhealthy.
- The service hit a resource limit and is load-shedding.
- A transient network partition.
When to retry: Always. Retry with exponential backoff. The service will likely recover.
DEADLINE_EXCEEDED (4)
The client specified a deadline (e.g., "I need an answer in under 500ms") and the server didn't meet it. This can happen because:
- The server is overloaded and can't process the request in time.
- The actual work takes longer than the client's deadline allows.
- There's network latency that pushes the total time over the limit.
When to retry: Only if you increase the deadline or accept that you'll hit it again. Blindly retrying the same deadline is futile and wastes resources. See deadline exceeded code 4 for depth.
ABORTED (10)
The operation was aborted due to a concurrency issue — typically a conflict that prevented the operation from completing atomically. Examples:
- A transaction conflict in a database.
- A check-and-set operation that failed because the value changed between check and set.
When to retry: Yes. Aborted is a signal that the operation could succeed if retried, because the conflicting state may change. Exponential backoff is wise.
FAILED_PRECONDITION (9)
The system is in the wrong state for this operation. Unlike ABORTED, this is not a transient conflict — it's a stable precondition violation. Examples:
- Trying to delete a non-empty bucket.
- Attempting an operation on a resource that's locked.
- Calling an RPC before authentication succeeds.
When to retry: Rarely. FAILED_PRECONDITION means you need to fix the precondition first (empty the bucket, unlock the resource, authenticate) before retrying. Blind retries won't help.
Official guidance: gRPC documentation explicitly groups UNAVAILABLE, ABORTED, and FAILED_PRECONDITION as the "retry family" — operations that may succeed on retry. But the why differs, and a smart retry strategy must account for that difference.
Other common codes and their causes
INTERNAL (13)
An unexpected error occurred inside the server — typically an unhandled exception or panic. This is the server saying "something broke, and I don't know how to describe it nicely." Examples:
- A null pointer dereference.
- An uncaught exception in handler code.
- A bug in the service's business logic.
Retrying: No. INTERNAL means the server encountered a bug, not a transient issue. Retrying will likely hit the same bug. Your monitoring system should alert on INTERNAL errors because they represent code defects.
UNKNOWN (2)
Something went wrong, and the server couldn't determine what. Often this is the result of an unhandled exception that bubbles up before the gRPC framework can wrap it in a proper status. In many cases, UNKNOWN is a sign of poor error handling — the server caught an exception but didn't map it to a specific gRPC status.
Retrying: No. It's genuinely unknown whether retrying will help. The root cause is server-side and requires investigation.
UNIMPLEMENTED (12)
The RPC method doesn't exist on the server. This usually means:
- The service doesn't have the method defined.
- The client is calling a newer API version that the server doesn't yet support.
- A typo in the method name.
Retrying: No. The method will never exist on this server, so retrying is pointless.
NOT_FOUND (5)
The requested resource does not exist. This is a successful operation — the server understood the request, but the resource isn't there. Examples:
GET /user/404returns NOT_FOUND, not an error.
Retrying: No. If the resource isn't there now, it won't be there in 100ms.
Who generates the status code?
Understanding the source of a status code helps you debug. Status codes come from three places:
- The client library — for errors it can detect before sending (invalid arguments, cancelled deadlines).
- The server — the RPC handler returns a status.
- The load balancer or proxy — for network-level issues (connection refused, timeout on the wire).
When the load balancer can't reach any backend, it returns UNAVAILABLE. When the server's handler panics, the framework catches it and returns INTERNAL. When the client detects its deadline has passed, it cancels the call and returns CANCELLED.
Inspecting status codes in code
Go
package main
import (
"context"
"log"
"google.golang.org/grpc"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
)
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 500*time.Millisecond)
defer cancel()
resp, err := client.GetUser(ctx, &pb.GetUserRequest{Id: "123"})
if err != nil {
st, ok := status.FromError(err)
if !ok {
log.Fatalf("error is not a gRPC status: %v", err)
}
log.Printf("Code: %v (%s)", st.Code(), st.Code().String())
log.Printf("Message: %s", st.Message())
// Check if retryable
switch st.Code() {
case codes.Unavailable, codes.ResourceExhausted, codes.Aborted:
log.Println("Retryable error; consider backing off and trying again")
case codes.DeadlineExceeded:
log.Println("Deadline exceeded; retrying won't help unless you increase the deadline")
default:
log.Println("Non-retryable error")
}
return
}
log.Printf("User: %v", resp)
}
Node.js
const grpc = require('@grpc/grpc-js');
const { status } = require('@grpc/grpc-js');
async function getUser() {
try {
const response = await client.getUser({ id: '123' });
console.log('User:', response);
} catch (error) {
// error is a RpcError
const code = error.code;
const message = error.message;
console.log(`Code: ${code} (${status[code]})`);
console.log(`Message: ${message}`);
if ([
status.UNAVAILABLE,
status.RESOURCE_EXHAUSTED,
status.ABORTED,
].includes(code)) {
console.log('Retryable; back off and retry');
} else if (code === status.DEADLINE_EXCEEDED) {
console.log('Deadline exceeded; increase it and try again, or fail');
} else {
console.log('Non-retryable error');
}
}
}
Java
import io.grpc.Status;
import io.grpc.StatusRuntimeException;
try {
GetUserRequest request = GetUserRequest.newBuilder().setId("123").build();
GetUserResponse response = stub.getUser(request);
System.out.println("User: " + response);
} catch (StatusRuntimeException e) {
Status status = e.getStatus();
System.out.println("Code: " + status.getCode());
System.out.println("Message: " + status.getDescription());
switch (status.getCode()) {
case UNAVAILABLE:
case RESOURCE_EXHAUSTED:
case ABORTED:
System.out.println("Retryable; back off and retry");
break;
case DEADLINE_EXCEEDED:
System.out.println("Deadline exceeded; retry with increased deadline or fail");
break;
default:
System.out.println("Non-retryable error");
}
}
Retry strategies and idempotency
Retrying is safe only if your operation is idempotent — calling it multiple times has the same effect as calling it once. Reading a value is always idempotent. Writing a value is idempotent only if the operation is designed that way (e.g., set a field, overwrite a record). Incrementing a counter is not idempotent.
Before you retry on UNAVAILABLE, ABORTED, or RESOURCE_EXHAUSTED, ask: Does my operation have side effects that can accumulate? If yes, you need an idempotency key (a unique identifier for the operation so the server can detect and skip duplicate executions).
gRPC doesn't mandate idempotency, so it's your responsibility. See gRPC error handling and monitoring for a deeper discussion.
Use exponential backoff with jitter when retrying. Start at 100ms, double each time, cap at 10s, and add randomness to prevent thundering herds of retries. A formula: delay = min(cap, base * 2^attempt) + random(0, jitter).
Reducing errors with distributed tracing
The best way to catch gRPC errors early is distributed tracing. By recording the full span of each RPC call — including the status code, latency, and upstream/downstream dependencies — you can spot patterns: Is UNAVAILABLE spiking? Is one backend consistently returning DEADLINE_EXCEEDED? Are certain methods frequently returning UNIMPLEMENTED?
Instrumenting your gRPC services with tracing (and pairing it with circuit-breaker patterns and tail-latency monitoring) transforms raw status codes into actionable insight.
Do not confuse gRPC status codes with HTTP status codes. A gRPC service returning UNAVAILABLE (code 14) is sent over HTTP as a 503 Service Unavailable, but the gRPC code is what matters to your client. HTTP-only monitoring may miss the distinction.
Implementing server-side status handling
When your service is the server, return the most specific status code you can. Throwing a generic INTERNAL error is lazy. Instead, validate inputs and return INVALID_ARGUMENT; check preconditions and return FAILED_PRECONDITION; detect conflicts and return ABORTED. This gives clients signal to make intelligent retry decisions.
If your service lacks a rich error model (error details attached to the status), at minimum include a descriptive message. "resource exhausted: connection pool at capacity" is more useful than "error."
Summary
gRPC's 17 status codes are a contract between client and server. Respecting that contract — using the right code in the right place, retrying intelligently, and monitoring for patterns — keeps your microservices reliable and your debugging tractable.
Start tracking errors in minutes
Monitor gRPC status codes and latency patterns across your services with distributed tracing to catch errors before they cascade — start free with LightTrace.