Проблема
В монолите сервисы живут в одном процессе. В микросервисах нужно ответить на вопрос: "Где сейчас живёт инстанс сервиса payment?". Ответ нетривиален:
- Инстансы появляются и умирают (auto-scaling, rolling deploy, node crash)
- IP-адреса меняются (особенно в Kubernetes и контейнерных платформах)
- Инстансов может быть 1, 10 или 100 -- надо балансировать
- Нездоровые инстансы (high latency, errors) должны автоматически исключаться
Service discovery -- это механизм, который даёт consumer'у актуальный список здоровых endpoint'ов для сервиса по логическому имени.
Две модели: client-side vs server-side
Client-side discovery
Клиент сам запрашивает реестр, получает список endpoint'ов и балансирует нагрузку сам.
[Client] --1. where is "payment"?--> [Registry (Consul/Eureka)]
[Client] <--2. [10.0.1.5, 10.0.1.6, 10.0.1.7]--
[Client] --3. pick one (round-robin/weighted)--> [Instance]
Плюсы: нет дополнительного hop, клиент знает про все инстансы, можно делать smart load balancing (weighted, outlier detection).
Минусы: каждый язык/фреймворк нуждается в своём discovery-клиенте; связность выше; client должен уметь health check.
Реализации: Netflix Eureka + Ribbon, Finagle, Consul с DNS/HTTP API.
Server-side discovery
Клиент вызывает единый endpoint балансировщика. Балансировщик (не сам клиент) знает, куда направить запрос.
[Client] --> [Load Balancer / Proxy] --> [Instance]
^
| reads from
[Registry (K8s API / cloud LB)]
Плюсы: клиент не знает про discovery -- просто DNS имя или IP балансировщика; язык-агностично; один компонент ответственен за health check.
Минусы: лишний hop = latency; балансировщик -- single point of failure (обычно HA); sticky session сложнее.
Реализации: Kubernetes Service (kube-proxy + iptables/IPVS), AWS ALB/NLB, Google Cloud Internal LB, HAProxy.
Сравнение
| Критерий | Client-side | Server-side |
|---|---|---|
| Latency | ниже (без hop) | +1 hop (0.1-1 ms) |
| Smart LB (weighted, EWMA) | да | ограничено LB |
| Зависимость от языка | SDK под каждый | нет |
| Сложность клиента | высокая | низкая |
| Failover клиента | сам реализует | LB делает |
DNS-based discovery (Kubernetes)
В Kubernetes это по умолчанию. Объект Service имеет стабильное DNS-имя, за ним стоит набор Pod'ов, kube-proxy маршрутизирует.
# payment-service.yaml
apiVersion: v1
kind: Service
metadata:
name: payment
namespace: prod
spec:
selector:
app: payment
ports:
- port: 8080
targetPort: 8080
# type: ClusterIP - default, virtual IP inside cluster
Клиенты обращаются к payment.prod.svc.cluster.local:8080. CoreDNS резолвит имя в ClusterIP, kube-proxy балансирует на здоровые Pod'ы (на основе readinessProbe).
Headless services
Если нужны прямые IP Pod'ов (например, для gRPC client-side balancing или stateful workload):
apiVersion: v1
kind: Service
metadata:
name: payment-headless
spec:
clusterIP: None # headless
selector:
app: payment
ports:
- port: 8080
DNS-запрос к payment-headless вернёт все IP Pod'ов (A-records). Клиент сам выбирает один -- это client-side discovery поверх DNS.
Плюсы K8s DNS
- Работает из коробки, любой язык
- Kubelet + readinessProbe автоматически исключает нездоровые Pod'ы
- Ничего специального в клиентах
Минусы
- DNS caching: стандартный net/http в Go кэширует DNS, при быстром скейлинге клиент может держать IP умерших Pod'ов
- Нет smart load balancing: kube-proxy делает round-robin на iptables/IPVS
- gRPC долго держит соединение -- новые Pod'ы не получают траффика пока не отключили старые
Go клиент с K8s DNS + connection pool
Базовый HTTP-клиент внутри кластера. Ключевые моменты: короткий IdleConnTimeout, разумный пул, отдельный Transport под каждый upstream.
package httpclient
import (
"context"
"fmt"
"net"
"net/http"
"time"
)
// NewPooledClient builds an HTTP client tuned for Kubernetes service discovery.
//
// Why these settings:
// - DialContext with short DNS cache: pick up new Pod IPs quickly
// - MaxIdleConnsPerHost: enough to reuse, not too many to hoard
// - IdleConnTimeout: drop idle conns to free old Pod IPs
// - ForceAttemptHTTP2: multiplex over single connection where possible
func NewPooledClient(timeout time.Duration) *http.Client {
dialer := &net.Dialer{
Timeout: 3 * time.Second,
KeepAlive: 30 * time.Second,
}
return &http.Client{
Timeout: timeout,
Transport: &http.Transport{
DialContext: dialer.DialContext,
MaxIdleConns: 200,
MaxIdleConnsPerHost: 32,
MaxConnsPerHost: 64,
IdleConnTimeout: 30 * time.Second, // drop stale conns to dead Pods
TLSHandshakeTimeout: 3 * time.Second,
ExpectContinueTimeout: 1 * time.Second,
ForceAttemptHTTP2: true,
ResponseHeaderTimeout: 5 * time.Second,
},
}
}
// PaymentClient talks to Kubernetes service "payment".
// DNS: payment.prod.svc.cluster.local resolves to ClusterIP.
type PaymentClient struct {
baseURL string // http://payment.prod:8080
http *http.Client
}
func NewPaymentClient(baseURL string) *PaymentClient {
return &PaymentClient{
baseURL: baseURL,
http: NewPooledClient(5 * time.Second),
}
}
func (c *PaymentClient) Charge(ctx context.Context, orderID string, amountCents int64) (string, error) {
req, err := http.NewRequestWithContext(ctx, "POST",
fmt.Sprintf("%s/charges", c.baseURL), nil)
if err != nil {
return "", err
}
req.Header.Set("Idempotency-Key", orderID)
resp, err := c.http.Do(req)
if err != nil {
return "", fmt.Errorf("http call: %w", err)
}
defer resp.Body.Close()
if resp.StatusCode >= 500 {
// 5xx = retryable, caller should use circuit breaker + retry
return "", fmt.Errorf("upstream 5xx: %d", resp.StatusCode)
}
// parse body...
return "pay_xxx", nil
}
gRPC client-side balancing через DNS headless
import (
"google.golang.org/grpc"
"google.golang.org/grpc/balancer/roundrobin"
_ "google.golang.org/grpc/resolver/dns"
)
// dns:///payment-headless.prod.svc.cluster.local:50051
// Scheme "dns" + triple slash uses built-in DNS resolver
// that returns all Pod IPs from headless service.
conn, err := grpc.NewClient(
"dns:///payment-headless.prod:50051",
grpc.WithDefaultServiceConfig(`{"loadBalancingConfig":[{"round_robin":{}}]}`),
grpc.WithTransportCredentials(insecure.NewCredentials()),
)
Здесь gRPC резолвит headless service периодически (default 30s), держит открытыми соединения ко всем Pod'ам, балансирует round-robin на уровне запросов (не соединений).
Consul service discovery
Consul -- популярный реестр вне Kubernetes (или совместно с ним). Сервисы регистрируются, Consul проверяет их health check, клиенты читают каталог по HTTP API или DNS.
Регистрация
{
"service": {
"name": "payment",
"id": "payment-1",
"address": "10.0.1.5",
"port": 8080,
"tags": ["v2", "primary"],
"check": {
"http": "http://10.0.1.5:8080/health",
"interval": "10s",
"timeout": "2s",
"deregister_critical_service_after": "1m"
}
}
}
Обычно регистрацию делает sidecar (consul-agent) или сам сервис на старте (PUT /v1/agent/service/register).
PHP Symfony клиент с Consul catalog
Читаем список здоровых инстансов из Consul, кэшируем на короткое время, балансируем round-robin. В продакшне лучше использовать sidecar, но для понимания:
<?php
declare(strict_types=1);
namespace App\Discovery;
use Symfony\Contracts\HttpClient\HttpClientInterface;
use Symfony\Contracts\Cache\CacheInterface;
use Symfony\Contracts\Cache\ItemInterface;
use Psr\Log\LoggerInterface;
/**
* Discovers healthy instances of a service via Consul HTTP API.
* Caches for short TTL (5s) to reduce load on Consul.
*/
final class ConsulDiscovery
{
public function __construct(
private readonly HttpClientInterface $consul, // base_uri: http://consul:8500
private readonly CacheInterface $cache,
private readonly LoggerInterface $logger,
) {}
/**
* @return list<array{address: string, port: int}>
*/
public function healthy(string $serviceName): array
{
return $this->cache->get(
"consul:$serviceName",
function (ItemInterface $item) use ($serviceName): array {
$item->expiresAfter(5); // short TTL - balance freshness and load
try {
$resp = $this->consul->request('GET',
"/v1/health/service/$serviceName",
['query' => ['passing' => 'true'], 'timeout' => 2.0]
);
$data = $resp->toArray();
} catch (\Throwable $e) {
$this->logger->error('consul unavailable', ['err' => $e->getMessage()]);
return [];
}
$out = [];
foreach ($data as $entry) {
$svc = $entry['Service'];
$out[] = [
'address' => $svc['Address'] ?: $entry['Node']['Address'],
'port' => $svc['Port'],
];
}
return $out;
}
);
}
}
/**
* Round-robin balancer over discovered instances.
* Uses APCu for cross-request counter in single process; for multi-process
* use Redis INCR or ETS-like store.
*/
final class RoundRobinBalancer
{
public function __construct(
private readonly ConsulDiscovery $discovery,
) {}
public function pick(string $service): ?string
{
$instances = $this->discovery->healthy($service);
if (empty($instances)) {
return null;
}
$key = "rr:$service";
$idx = apcu_inc($key, 1);
$chosen = $instances[$idx % count($instances)];
return sprintf('http://%s:%d', $chosen['address'], $chosen['port']);
}
}
/**
* Thin payment client using discovery.
* Retries on transport error with fresh instance selection (outlier ejection).
*/
final class PaymentClient
{
private const MAX_ATTEMPTS = 3;
public function __construct(
private readonly HttpClientInterface $http,
private readonly RoundRobinBalancer $balancer,
private readonly LoggerInterface $logger,
) {}
public function charge(string $orderId, int $amountCents): string
{
$lastError = null;
for ($attempt = 1; $attempt <= self::MAX_ATTEMPTS; $attempt++) {
$base = $this->balancer->pick('payment');
if ($base === null) {
throw new \RuntimeException('no healthy payment instances');
}
try {
$resp = $this->http->request('POST', "$base/charges", [
'json' => ['order_id' => $orderId, 'amount_cents' => $amountCents],
'headers' => ['Idempotency-Key' => $orderId],
'timeout' => 5.0,
]);
return $resp->toArray()['payment_id'];
} catch (\Throwable $e) {
$this->logger->warning('payment call failed', [
'attempt' => $attempt,
'base' => $base,
'err' => $e->getMessage(),
]);
$lastError = $e;
// Next iteration - another instance via round-robin
usleep(100_000 * $attempt); // linear backoff
}
}
throw new \RuntimeException('payment unavailable', 0, $lastError);
}
}
DNS-интерфейс Consul
Consul также экспонирует DNS: payment.service.consul резолвится в A-records здоровых инстансов. Удобно, если SDK нет -- просто HTTP-клиент с DNS именем.
Service mesh dataplane: Envoy
Service mesh (Istio, Linkerd, Consul Connect) реализует discovery через sidecar-прокси. Клиент всегда обращается к localhost -- sidecar (обычно Envoy) знает всё остальное.
[App] --localhost:9000--> [Envoy sidecar] --mTLS--> [Envoy sidecar] --localhost--> [App]
|
config from control plane (Istiod / Consul)
(EDS - Endpoint Discovery Service)
Плюсы: язык-агностично, mTLS, ретраи, circuit breaker, outlier detection, observability -- всё на уровне mesh. Приложение просто делает HTTP к 127.0.0.1. Минусы: оверхед (CPU/memory на sidecar), сложность, extra latency ~1-3 ms.
Подробнее в 6.service-mesh.md.
Health check integration
Discovery без health check -- ложная безопасность. Типы проверок:
| Тип | Что проверяет | Для чего |
|---|---|---|
| Liveness | Процесс жив | Перезапустить Pod/контейнер |
| Readiness | Готов принимать трафик | Исключить из LB пока не готов |
| Startup | Начальная готовность | Долгий старт (JVM, миграции) |
Важно: readinessProbe и liveness probe в Kubernetes -- это разные вещи. Readiness управляет маршрутизацией (исключение из Service endpoints), liveness -- рестартом.
Пример в Go
// /health/ready checks external deps; /health/live only checks that process runs.
http.HandleFunc("/health/live", func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusOK)
})
http.HandleFunc("/health/ready", func(w http.ResponseWriter, r *http.Request) {
ctx, cancel := context.WithTimeout(r.Context(), 500*time.Millisecond)
defer cancel()
if err := db.PingContext(ctx); err != nil {
http.Error(w, "db down", http.StatusServiceUnavailable)
return
}
if err := cache.Ping(ctx).Err(); err != nil {
http.Error(w, "cache down", http.StatusServiceUnavailable)
return
}
w.WriteHeader(http.StatusOK)
})
# Kubernetes probes
readinessProbe:
httpGet: { path: /health/ready, port: 8080 }
periodSeconds: 5
failureThreshold: 3
livenessProbe:
httpGet: { path: /health/live, port: 8080 }
periodSeconds: 10
failureThreshold: 5
initialDelaySeconds: 30
Graceful shutdown
При получении SIGTERM: сначала убрать себя из discovery (failing readiness), подождать ~10 сек (чтобы LB перестал слать трафик), потом закрыть соединения.
srv := &http.Server{Addr: ":8080", Handler: mux}
go srv.ListenAndServe()
<-signalCh // SIGTERM
atomic.StoreInt32(&ready, 0) // readiness now fails
time.Sleep(10 * time.Second) // drain
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
srv.Shutdown(ctx)
Сравнительная таблица реализаций
| Инструмент | Тип | Где используется | Плюс | Минус |
|---|---|---|---|---|
| Kubernetes Service | Server-side (DNS + iptables) | Внутри K8s | Из коробки | Простой LB, gRPC проблемы |
| Kubernetes Headless | Client-side (DNS) | gRPC внутри K8s | Прямые IP | Нет health за пределами readiness |
| Consul | Both | Hybrid, VM + K8s | Гибкость, DC-aware | Ещё один компонент |
| Eureka (Netflix) | Client-side | JVM экосистема | Проверено годами | Устаревает |
| AWS ALB + ECS | Server-side | AWS managed | Нулевая настройка | Vendor lock |
| Envoy (Istio) | Dataplane | Service mesh | Smart LB, mTLS | Сложность, оверхед |
Типичные ошибки
- Долгое DNS-кэширование -- Go
net/httpне перевычитывает DNS в keep-alive соединениях, старые Pod IP удерживаются часами. Лечение: короткийIdleConnTimeout, health check на клиенте, или HTTP/2 без keep-alive. - Нет outlier detection -- один медленный инстанс тянет весь round-robin. Решение: EWMA-балансировка или circuit breaker per-instance.
- Regist -> crash -> zombie -- сервис зарегистрировался, упал не по SIGTERM, остался в реестре. Спасает только TTL/health check:
deregister_critical_service_afterв Consul, или агент-based регистрация. - Readiness = liveness -- частая ошибка: endpoint проверяет БД и используется и там, и там. В итоге при краткой недоступности БД Kubernetes рестартует все Pod'ы.
- Discovery без retry -- получил IP, получил connection refused, вернул ошибку пользователю. Нужен retry на свежем инстансе.
Выводы
Выбор модели service discovery зависит от платформы: в Kubernetes стандарт -- DNS-based ClusterIP (server-side); для gRPC и продвинутой балансировки -- headless + client-side; в гибридных средах -- Consul; в service mesh -- Envoy sidecars. В любой реализации критичны три вещи: health check, который реально проверяет готовность; graceful shutdown, снимающий readiness до закрытия соединений; retry на клиенте при transport-ошибках с перевыбором инстанса.