DevOps&SRE Library


Гео и язык канала: Россия, Русский
Категория: Технологии


Библиотека статей по теме DevOps и SRE.
Реклама: @ostinostin
Контент: @mxssl
РКН: https://www.gosuslugi.ru/snet/67704b536aa9672b963777b3

Зарегистрирован в РКН
Связанные каналы  |  Похожие каналы

Гео и язык канала
Россия, Русский
Категория
Технологии
Статистика
Фильтр публикаций


Benchmarking LLM Inference with Production Agent Traces

Most load-testing tools, however, were designed for a simpler world. They assume requests are independent, stateless, and interchangeable. But an AI agent doesn't work that way. This gap was the motivation behind OpenTelemetry (OTel) Trace Replay, a new capability in Inference Perf, the Kubernetes SIG tool for GenAI inference benchmarking.


https://medium.com/inference-perf/benchmarking-llm-inference-with-production-agent-traces-f47f7f994aff


Cordium - Kubernetes sandboxes with secretless access

Cordium is a free and open source, self-hosted, identity-based sandbox platform built on Kubernetes and Octelium. Cordium is a general-purpose platform that provides isolated, reproducible isolated sandboxes for developers, AI agents, and automated workloads.


https://github.com/octelium/cordium


Ballast: Kubernetes right-sizing operator

Ballast is a Kubernetes operator that automatically right-sizes workload resource requests and limits based on real operational history. It is a more active alternative to Fairwinds Goldilocks: rather than suggesting changes, it applies them — at admission time and on running pods via in-place resize (Kubernetes 1.35+).


https://github.com/Tight-Line/ballast


Этот пост видят только приглашенные на конференцию

21 октября для вас и ваших коллег Т-Банк проводит онлайн-встречу «SRE-техтолк». Приглашают SRE, DevOps-инженеров и разработчиков.

Практикующие эксперты приготовили классную программу с реальными кейсами и без абстракций. Вас ждет:
Разбор задач Т-Банка и эффективных практик.
Подходы к надежности, о которых редко рассказывают публично.
Адаптация инструментов под разные задачи.

Приходите послушать трех специалистов с разносторонним опытом, обсудить ваши задачи и погрузиться в бигтех.

Встреча пройдет онлайн. Участие бесплатное, а регистрация — по ссылке!


ArgoCD Mobile

A native iOS and Android app for monitoring and managing Argo CD deployments from your phone.


https://github.com/argoproj-labs/mobile-for-argocd


Заглянем внутрь современных инженерных платформ

14 октября Т-Банк приглашает на Platform Engineering Night — вечер для инженеров, которые создают платформы и помогают командам справляться с растущей сложностью разработки.

Что будем делать:
🔴Обсудим внутренние платформы, которые помогают быстрее разрабатывать, запускать и сопровождать сервисы.
Разберем, как масштабировать платформы, когда растет количество пользователей, сервисов и связей между ними.
🔴Расскажем про новые Т-ЦОДы — обсудим, как платформы помогают согласовать работу между дата-центрами в разных регионах.
🔴Покажем, как AI встроен в платформы и весь SDLC — от разработки до эксплуатации.
🔴Обсудим реальные инженерные кейсы с командами Т-Банка и других технологических компаний.

➡Когда: 14 октября, начало в 18:00
📍Где: Онлайн и офлайн в ИТ-хабе Т-Банка в Санкт-Петербурге

Мероприятие бесплатное, торопитесь занять место по ссылке на регистрацию.


l9gpu: GPU telemetry with workload attribution

DCGM exporter tells you a GPU is hot. It won't tell you whose job is frying it. l9gpu closes the loop. One agent per node emits vendor-neutral OTLP with workload attribution baked in — Kubernetes pod, namespace, deployment; Slurm job, user, partition.


https://github.com/last9/gpu-telemetry


Authentication between microservices using Kubernetes identities

In this article, we will use Kubernetes Service Accounts and the TokenReview API to authenticate requests between two in-cluster services, then improve the setup with audience-bound projected Service Account tokens.


https://learnkube.com/microservices-authentication-kubernetes


Architecting Next-Generation Infrastructure: Enterprise Multi-Cluster Management Using Rancher, Virtual Clusters (vCluster), and Kargo GitOps

This guide breaks down the ultimate architectural blueprint to solve cluster sprawl and deployment friction for good. By unifying the centralized governance of Rancher, the resource-saving isolation of virtual control planes (vCluster), and the advanced multi-stage pipeline automation of Kargo GitOps, this write-up touches on how to build a highly scalable, zero-touch infrastructure environment.


https://medium.com/@taofeekaoyusuf/architecting-next-generation-infrastructure-enterprise-multi-cluster-management-using-rancher-fd024a400c87


Building a Kubernetes Raspberry Pi Homelab with K3s

As you'll know if you've ever tried to build a reasonably sized microservices application locally, you can run out of system resources quickly. So, in order to free up some of that valuable memory, I figured I'd offload things to a Kubernetes cluster running on some relatively cheap hardware, and in this post I'll describe the steps I took so that you can do the same.


https://medium.com/@chris.allmark/building-a-kubernetes-raspberry-pi-homelab-with-k3s-f934e4b24162


The kubectl 401 That Wasn't a Kubernetes Problem

A short debugging story about a two-year-old credentials file, an EKS cluster, and the AWS credential chain quietly doing exactly what it was designed to do.


https://medium.com/@pradhyuman-pandey/the-kubectl-401-that-wasnt-a-kubernetes-problem-4e04438f33ae


Securing CI/CD for an open source project: lessons from Cilium

Cilium runs in the kernel-level networking path of millions of Kubernetes pods. If our supply chain were compromised, the blast radius would not be small. Hardening the project against that scenario is something we work on continuously, and we wanted to write down what we actually do, in detail. Most of what follows isn't Cilium-specific: any open source project running CI/CD on GitHub Actions can apply these patterns.


https://cilium.io/blog/2026/05/06/securing-cicd-open-source-lessons-from-cilium


How I Learned to Stop Worrying and Love the Reconciliation Loop

You've got ten Claude Code agents running. Each is working on a different issue. Agent 3 just created a PR that conflicts with Agent 7's changes. Agent 5 has been stuck on "Analyzing codebase…" for 45 minutes. Welcome to multi-agent chaos. Turns out, scaling AI coding agents from 1 to 10+ is a completely different problem from running a single agent in a terminal.


https://medium.com/@bobbydeveaux/how-i-learned-to-stop-worrying-and-love-the-reconciliation-loop-32928d0d80cf


Миссия на сегодня:
— обновить платформу, не обновляя весь Kubernetes;
— встроить AppSec, не остановив релизы;
— пережить падение хранилища секретов;
— защитить LLM и понять, куда исчезли все токены;
— усложнить Admission-политики, не положив API-сервер.

Если звучит как особенно бодрое дежурство — приходите 15 октября на DevOps-трек конференции Orion soft «Большая игра». Там эти сценарии разберут по архитектуре, коду и результатам тестов.

📍 Москва
Принять миссию


Kubernetes is migrating from SPDY to WebSockets

I maintain an app that builds on top of kubernetes port forwarding, so i track KEP-4006 because the streaming protocol underneath keeps changing. I wrote about this back in april 2024 around Kubernetes 1.30, and six releases later it's changed enough to be worth another look.


https://kftray.app/blog/kubernetes-spdy-to-websockets


Server-side apply: what happens when you run kubectl apply

Server-side apply matters because Kubernetes objects are shared state: it moves field ownership into the API server, so apply-style tools can surface conflicts instead of hiding them as silent overwrites.


https://learnkube.com/server-side-apply-kubernetes


How Netflix Simplified Batch Compute with Kueue

As a part of the journey to transition Netflix's compute infrastructure to be more Kubernetes-native, we have leaned into incorporating components from the Kubernetes ecosystem into our container platform Titus. One example of this is our use of Kueue, a cloud-native job queueing system for batch workloads, which has largely replaced the custom queuing and scheduling logic in our homegrown managed batch solution Compute Managed Batch (CMB).


https://medium.com/netflix-techblog/how-netflix-simplified-batch-compute-with-kueue-87860682629c


openrig

A harness wraps a model. A rig wraps your harnesses. Define your agent team in YAML, boot it with one command. Claude Code and Codex in the same rig, managed as one system.

OpenRig turns AI coding agents from a pile of terminal sessions into a persistent, organized team. Talk to a lead agent about the outcome you want; it can coordinate specialists across teams and bring you results and decisions that need your attention. Start with a repository and one useful change, then keep the team's work and context at the same addresses.


https://github.com/mvschwarz/openrig


whiteboard

Whiteboard is an open-source desktop app where humans and agents can architect software together in a common workspace.


https://github.com/devdotfast/whiteboard


walgithub

walgit with a GitHub Enterprise Server facade — git over smart HTTP from an object-store bucket


https://github.com/rgodha24/walgithub

Показано 20 последних публикаций.