In Kubernetes, resource allocation has historically been a static decision made during a Pod's initial scheduling and placement. With the graduation of the core in-Place Pod resize feature to…
Kubernetes Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Labels are a powerful way to organize telemetry and define policies across Grafana Cloud, helping to streamline alerting, attribution, access control, and more. But traditionally, custom labels in…
Grafana Labs
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Scope This document describes three failure scenarios that separate having backups from being able to recover, and the guidance that follows from each. Every scenario is reproducible on a laptop…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Kubernetes has many ways to describe what is happening on a Node. Readiness, taints, Pod state, labels, annotations, and provider-specific APIs each expose part of the picture. What has been missing…
Kubernetes Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
If your Cypress suite has tests that fail more often or run slower, you know it can be hard to figure out the pattern from a single job. It could be one spec that slowed down, or a single test that…
Grafana Labs
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
The question that stopped the meeting It was a routine cost review. The slide showed the month’s GPU spend, the biggest line on the whole infrastructure bill, and someone asked a five-word question:…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
“A sales guy writing code” used to be the lead-up to a joke. But now no one’s laughing. Designers used to sit meekly waiting for the high priests of code to make their designs real. Now...
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
AI/ML and complex batch workloads continue to push the boundaries of Kubernetes scheduling. Following the foundational workload-centric enhancements introduced in previous releases, Kubernetes v1.37…
Kubernetes Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Access control belongs on the same day-zero checklist as networking and storage. On most on-prem clusters, it never makes the list. The Identity Gap Managed cloud Kubernetes ships IAM or SSO…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
You’ve probably felt this one: GitHub Actions usage creeps up across your org, and your actual visibility into it doesn’t keep pace. Which workflows are slow? Which are flaky? How long are jobs…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Multi-cluster, multi-cloud Kubernetes orchestration project reaches production maturity as global enterprises scale AI training and inference across hybrid infrastructure Key Highlights SHANGHAI,…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
New members including SoftBank Corp. and Crusoe join the cloud native community to help build cost-efficient, sovereign infrastructure SHANGHAI, China – KubeCon + CloudNativeCon + OpenInfra Summit +…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
New cloud native platform lifted average accelerator compute utilization from 35% to more than 60% and cut inference cost per 1 million tokens by more than 60% Key Highlights SHANGHAI, China –…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
New research finds China’s IIoT developers (48%) outpace the global average (42%) in cloud native adoption as AI infrastructure matures Key Highlights: SHANGHAI – KubeCon + CloudNativeCon +…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Recipe Handling Vulnerability Reports Target audience (the chef) This recipe is aimed at small and medium non-security focused projects. Maintainers of a high-risk security sensitive project, you…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Kubernetes v1.37 promotes the KubeletInUserNamespace feature gate to beta. With this feature enabled, all of the node components (kubelet, CRI and OCI runtimes, CNI plugins, and kube-proxy) can run…
Kubernetes Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Migrating a production database is a high-risk operational event. AWS DMS is a cloud service that migrates relational databases, data warehouses, and other data stores into the AWS Cloud or between…
AWS DevOps Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
AI infrastructure conversations often start with GPUs. Accelerators provide much of the compute behind model training and inference, so the focus is understandable. But a production AI workload…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Kubernetes isn’t brand new anymore. Yet, for many teams, adopting it still feels intimidating. Even if you’ve watched Kubernetes become the default foundation for production software and AI…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
A metrics dashboard can tell you a system’s health with ease. A log can help you understand a discrete failure. The post How to find failures without drowning in tracing data appeared first on The…
The New Stack
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Kubernetes 1.37 is here and Dynamic Resource Allocation (DRA) keeps pushing past where it started! This release brings DRA Extended Resource support to GA, a milestone the team has been building…
Kubernetes Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
There’s still time to join OSPOlogy + OSPO Summit China 2026, taking place on September 7, 2026, in Shanghai, China as part of KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China.…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Somewhere in your cluster there’s probably a deployment sitting in the default namespace that everyone knows shouldn’t be there. Nobody put it there maliciously, it just happened, early on, before…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Introduction Continuous improvement depends on experimentation. Teams know that the fastest path to better outcomes is to test changes against real user behavior, measure results, and iterate. In…
AWS DevOps Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Anthropic aimed to steer its ship into safer, more carefully charted waters this week. The company announced it was improving The post Anthropic’s Claude failures have made agent observability a…
The New Stack
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
AWS Health Planned Lifecycle Events signal when a managed service version is nearing end of standard support. Learn how to automate these upgrades end to end with AWS DevOps Agent and Kiro: detect…
AWS DevOps Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Kubernetes v1.37 includes API support for horizontal autoscaling of workloads down to zero replicas. This feature is now Beta and enabled by default. A HorizontalPodAutoscaler (HPA) that uses a…
Kubernetes Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
In the previous post , we introduced KubeVirtBMC and showed how it provides virtual BMC endpoints for KubeVirt VMs. We tested it with raw IPMI and Redfish commands. That was fun, but the real power…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
I am excited to announce that etcd RangeStream is graduating to beta in Kubernetes v1.37. Paired with etcd v3.7, it reduces the memory the API server and etcd need to read a large collection, and…
Kubernetes Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Most platform engineering conversations tend to split into two rooms pretty quickly. The first room is full of teams who don’t have a platform yet. Scattered scripts, tribal knowledge, and every…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
I am excited that storage version migration (SVM) has graduated to General Availability (GA) in Kubernetes v1.37! After a number of releases of work and testing, the built-in StorageVersionMigration…
Kubernetes Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
In case you missed it: OpenTelemetry (OTel) has officially achieved CNCF graduated status! It now stands proudly alongside amazing open source projects such as Kubernetes and Prometheus, to name…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Kubernetes made infrastructure more programmable, scalable, and resilient. It also made production systems harder to reason about. Workloads move, replicas churn, dependencies multiply, and a single…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Introduction Running workloads on Amazon Elastic Kubernetes Service (Amazon EKS) can involve managing failures like OOMKilled or IP exhaustion. Engineers must repeatedly collect pod logs, trace…
AWS DevOps Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Engineering organizations are trying to deliver as fast as technology allows, bringing agentic AI into their developer platforms and working The post The 3 roles AI agents play in your developer…
The New Stack
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Pod Certificate / Cluster Trust Bundles Blog Post Kubernetes brings a wealth of features that make it easy to run your production workloads securely and reliably. While aspects like scheduling,…
Kubernetes Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
The 3 AM Call We got paged one Tuesday morning. A critical production service had crashed under traffic—not gradually degraded, but crashed. Hundreds of pending pods. Users were seeing 15–20% error…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Kubernetes has given platform teams a consistent way to deploy, scale, and operate containerized applications. Now, many of those same teams are being asked to support AI. The transition is already…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Kubernetes v1.37 promotes the metrics.k8s.io API to stable ( v1 ). This API provides CPU and memory usage for nodes and Pods, and is the API behind commands such as kubectl top and…
Kubernetes Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
HCP Vault Dedicated has no native Microsoft Sentinel connector. Deploy a Terraform-managed pipeline that streams audit logs into Azure Log Analytics and Microsoft Sentinel.
HashiCorp
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Modern engineering teams instrument everything, with metrics, logs, traces, and profiles flowing from hundreds of services at once. But full-stack observability isn’t really about collecting more…
Grafana Labs
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
An AI factory is not just a model or a cluster. It is a pool of GPUs that many teams draw from at once: one team fine-tuning, another serving inference, a third running evaluations, all on...
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Observability is entering a new phase now that OpenTelemetry has standardized instrumentation for data collection. Unfortunately, the observability industry still The post Observability has a data…
The New Stack
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
This post was co-written with Michael Stephan, Senior Principal Product Manager, and Christian Kreuzberger, Principal Software Engineer, at Dynatrace. AI-driven software delivery changes how code…
AWS DevOps Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Clear patterns have emerged from governance reviews across 72 CNCF projects, distinguishing between what the CNCF requires at each maturity level versus what the data recommends for long-term…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Editors: Arsh Sharma, Christopher Tineo, Kirti Goyal, Sophia Ugochukwu, Swathi Rao, Troy Connor Similar to previous releases, the release of Kubernetes v1.37 introduces new Stable, Beta, and Alpha…
Kubernetes Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
Starting an experiment with an LLM has never been easier. Keeping a growing collection of those experiments consistent is another matter. Earlier this year, as more teams began exploring AI features…
Grafana Labs
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
It’s no secret that developers are increasingly being asked to shift left. It seems there’s always something new to shift left on. And now developers are being asked to shift left on observability.…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
As a recovering VMware architect, it took me a little while to grasp Kubernetes. And I noticed I’m not alone in this.. From developers on our own team who need to get fluent in Kubernetes fast...
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
The problem: Humans shouldn’t be correlation engines At Atlassian’s scale, hundreds of interconnected microservices distributed across multiple regions mean a production incident generates an…
CNCF Blog
Signal·Strong signalFire·This is fireBig Brain·Big brain moveShip It·Ship it!Noise·Just noise
CommentsSave for laterShare
Every LLM request splits into prefill (compute-bound, sets TTFT) and decode (memory-bound, sets TPOT). Understanding both is how you control latency and cost.