Table of Contents
- At a Glance: Why We Decided Against Cilium (TL;DR)
- 1. Why We Chose Cilium: The eBPF Vision
- 2. The Operational Reality: The Hidden Cost of eBPF
- 3. Direct CNI Comparison: Cilium vs. AWS VPC CNI vs. Calico
- 4. The Litmus Test: When Is Cilium Worth It – and When Is It Overengineering?
- 5. Why "Boring Infrastructure" Wins in Production
- 6. Our Way Back: How We Uninstalled and Migrated Away From Cilium
- 7. Conclusion: Choosing the Right Complexity for Your Team
- 8. Frequently Asked Questions (FAQ)
Meet the Author
2026-09-30
Field Report: Kubernetes Networking Without OverengineeringCilium: Time to say goodbye – Why we removed eBPF from our Kubernetes cluster
Cilium is widely considered a modern, high-performance solution for Kubernetes networking—and rightfully so. But not every cluster benefits from the additional complexity of eBPF, Hubble, and deep kernel interventions. In this field report, we share our real-world learnings with Cilium in a production AWS EKS cluster, why we eventually pulled the plug, and why "boring" network setups are often the more economical and resilient choice for many teams.

At a Glance: Why We Decided Against Cilium (TL;DR)
For anyone evaluating the right Container Network Interface (CNI) or considering introducing Cilium, here is a summary of our primary takeaways:
- Initial driver: We wanted to observe HTTP responses and status codes from outbound requests to external APIs via Hubble without invasive sidecars.
- Operational reality: Debugging shifted from proven Linux tools (
tcpdump,iptables) to specialized eBPF tools (cilium-dbg, BPF maps). This drastically increased the team's cognitive load. - Upgrades as a risk factor: Every Kubernetes and node AMI upgrade on AWS EKS required careful verification of eBPF compatibility and Cilium CRD versions.
- Resource overhead vs. added value: Running Cilium agents, Hubble Relay, and Hubble UI consumed measurable node resources for features we barely utilized outside of API monitoring.
- The consequence: Returning to the native AWS VPC CNI in combination with application-level monitoring via OpenTelemetry and standard Prometheus metrics.
- Conclusion: For hyperscale environments and complex service meshes, Cilium is outstanding. For standard web applications and medium-sized clusters, it is often textbook overengineering.
1. Why We Chose Cilium: The eBPF Vision
Anyone exploring modern Kubernetes networking in the cloud-native ecosystem inevitably encounters Cilium. Originated by Isovalent and now a graduated CNCF project, Cilium promises a revolution in the Linux kernel: Instead of routing packets through slow iptables chains or IPVS, Cilium hooks directly into the kernel's data path using eBPF (Extended Berkeley Packet Filter).
In our earlier foundational guide, Using Cilium for Kubernetes Networking and Observability, we explored these advantages in depth:
- High-performance routing: Near zero-latency packet processing at the kernel level.
- Advanced network policies: Layer-7 filtering (e.g., based on HTTP methods, paths, or DNS names).
- Integrated observability with Hubble: Transparency into every single flow across the cluster without code changes.
- Replacement for kube-proxy: Elimination of scalability limits in large
iptablesrule sets.
Our Specific Use Case: External API Observability
The final impetus for adopting Cilium in our AWS Kubernetes cluster (AWS EKS) was an ostensibly simple operational need: Our microservices communicate heavily with external third-party APIs (payment gateways, CRM systems, logistics providers).
We wanted to answer three key questions:
- How many external requests fail with HTTP status codes such as
429 Too Many Requestsor502 Bad Gateway? - What are the exact latencies to specific external FQDNs?
- Which pods are responsible for unexpected egress traffic?
The vision with Cilium and Hubble seemed ideal: No tedious integration of logging libraries into every single microservice, and no extra sidecar proxies like Envoy per pod. Hubble promised to capture all these Layer-7 metrics transparently from the kernel and expose them via the Hubble UI and Prometheus.
2. The Operational Reality: The Hidden Cost of eBPF
Initial implementation worked as advertised. Hubble delivered attractive service maps and granular flow logs. However, during several months of continuous production operations, the flip side of the coin emerged: Complexity doesn't vanish—it merely shifts.
A New Debugging Paradigm: When tcpdump and iptables No Longer Suffice
In a standard Linux and Kubernetes setup, engineers rely on a well-known, battle-tested toolset during network incidents: tcpdump, ping, traceroute, netstat, iptables-save, or conntrack. These utilities have been industry standards for decades and are second nature to system administrators.
Under Cilium, the playing field changes fundamentally:
- Because eBPF programs hook into TC (Traffic Control) and XDP layers, running
tcpdumpon the host interface often fails to capture the packets you expect. - Routing decisions and NAT translations happen within BPF maps residing in kernel memory.
- When packets are dropped, tailing standard system logs is no longer sufficient. Instead, you must invoke specialized tools like
cilium-dbg monitor --type dropor inspect maps viabpftool map dump.
For an on-call team, this means that when an obscure connectivity issue arises between two services at 3:00 AM, years of accumulated Linux networking expertise only go so far. Suddenly, on-call engineers must navigate BPF bytecode, tail calls, and map capacity limits. The cognitive load on the team increased dramatically.
# Before Cilium: Familiar packet tracing on the node
sudo tcpdump -nn -i any port 443
# With Cilium: Specialized eBPF commands and BPF map inspection
cilium-dbg monitor --type drop
cilium-dbg bpf nat list
cilium-dbg endpoint list
Upgrade Complexity: Dual Dependencies During Cluster Updates
In a managed cloud environment like AWS EKS, teams typically enjoy smooth, automated updates: AWS releases new node AMIs, worker nodes are rolled over, and workloads continue seamlessly.
With Cilium, this lifecycle became noticeably riskier:
- Kernel compatibility: Cilium depends heavily on specific kernel versions and BPF features provided by the host OS. A routine AMI update containing a minor kernel patch previously triggered unforeseen incompatibilities.
- Coupled Helm and CRD upgrades: Prior to every Kubernetes version upgrade, we had to coordinate Cilium Helm charts and their Custom Resource Definitions (
CiliumNetworkPolicy,CiliumClusterwideNetworkPolicy,CiliumNode). - Rollout troubleshooting: If connectivity issues occurred after a node restart, identifying the root cause was ambiguous: Was it AWS EKS, node security groups, or Cilium's eBPF bootstrap sequence?
Observability Overhead: Hubble for HTTP Monitoring?
Our primary motivation for adopting Cilium was observing external HTTP calls. After a few months, we conducted a realistic appraisal:
- Resource footprint: The
cilium-agenton each node along with central components likehubble-relayandhubble-uicontinuously consumed CPU and memory that was no longer available to user workloads. - High cardinality: Ingesting every single network flow generated immense data volumes that overburdened our Prometheus instances.
- Application-centric alternatives: We realized that standard application tracing solutions (such as OpenTelemetry or structured logs with Prometheus metrics in application code) addressed our requirement far better—with full context. The application knows not just that an API call returned 500, but also for which tenant and with which payload.
Hubble provided low-level kernel metrics, but business-level context was absent.
Documentation Gaps in Cloud Provider Integrations
While Cilium has extensive documentation, combining it with cloud-specific capabilities—such as AWS IAM Roles for Service Accounts (IRSA), AWS Security Groups per Pod, native ENI allocation, or AWS Load Balancer Controller—frequently revealed edge cases. Most online tutorials and community threads focus on bare-metal setups. Debugging hybrid architectures (Cilium chaining vs. Cilium native routing) demanded substantial investigation time.
3. Direct CNI Comparison: Cilium vs. AWS VPC CNI vs. Calico
To make an objective decision moving forward, we benchmarked the primary networking alternatives for AWS EKS:
| Evaluation Criterion | Cilium (eBPF) | AWS VPC CNI (Native) | Project Calico (iptables/BPF) |
|---|---|---|---|
| Architecture | Kernel eBPF data path | Native AWS ENI / IP management | Standard iptables / IPVS (optional eBPF) |
| AWS Integration | Third-party layer (overlay or ENI-IPAM) | 100% natively integrated into AWS VPC | Overlay (VXLAN/IPIP) or host routing |
| Performance & Throughput | Extremely high (minimal overhead) | Very high (direct VPC routability without NAT) | High to very high |
| Debugging Tools | cilium-dbg, Hubble, BPF maps | tcpdump, VPC Flow Logs, CloudWatch | iptables, calicoctl, standard Linux tools |
| Upgrade Effort | High (strict kernel/CRD dependencies) | Minimal (standard EKS managed add-on) | Moderate |
| Layer-7 Policies | Built-in (HTTP, Kafka, DNS) | Not native (requires service mesh/Calico) | Limited (primarily L3/L4) |
| Resource Overhead | Moderate to high (DaemonSet + Hubble) | Low (lightweight node daemon) | Low to moderate |
| Team Learning Curve | Steep (requires eBPF & BPF expertise) | Gentle (standard AWS networking concepts) | Gentle to moderate |
The conclusion was straightforward: For AWS EKS workloads, the AWS VPC CNI offers the simplest, most deeply integrated, and lowest-maintenance foundation. Every pod receives a real IP address from the AWS subnet, AWS VPC Flow Logs function natively, and AWS Security Groups can be assigned directly to pods.
4. The Litmus Test: When Is Cilium Worth It – and When Is It Overengineering?
Cilium is an outstanding project—in fact, it is one of the most technically ambitious open-source developments in cloud-native history. Yet, teams frequently fall into the trap of adopting solutions for architectural problems they do not actually experience.
Use the following decision tree to assess whether Cilium is warranted for your infrastructure:
┌────────────────────────────────────────┐
│ Does your cluster require a multi- │
│ cloud mesh or 5,000+ concurrent pods? │
└──────────────────┬─────────────────────┘
│
┌───────────────┴───────────────┐
YES NO
│ │
┌─────────▼────────┐ ┌───────────▼─────────────┐
│ Cilium │ │ Do you strictly require │
│ recommended │ │ kernel-level Layer-7 │
└──────────────────┘ │ network policies? │
└───────────┬─────────────┘
│
┌───────────────┴───────────────┐
YES NO
│ │
┌─────────▼────────┐ ┌───────────▼─────────────┐
│ Evaluate Cilium │ │ Stick with cloud-native │
└──────────────────┘ │ CNI (e.g. AWS VPC CNI / │
│ Azure CNI)! │
└─────────────────────────┘
Cilium is the right choice when...
- Massive scale is required: Your cluster encompasses hundreds of nodes and tens of thousands of pods, where
kube-proxyandiptablesreach their performance limits. - Cluster mesh / multi-cloud is mandatory: You connect multiple Kubernetes clusters across cloud providers transparently at the network layer.
- Strict zero-trust L7 policies are enforced: You must enforce granular HTTP and DNS policies without maintaining a dedicated service mesh like Istio or Linkerd.
- Dedicated eBPF expertise exists in-house: Your team has the skills to read kernel dumps and confidently debug BPF programs.
A simpler CNI is the better choice when...
- Small to medium clusters are operated: Workloads consist of typical web APIs, e-commerce platforms, or enterprise business software.
- Low maintenance is a top priority: Upgrades must proceed without auditing Helm release notes for breaking changes across BPF maps.
- Standard monitoring is sufficient: Application Performance Monitoring (APM), Prometheus, and Grafana fulfill your observability requirements.
- Standard Linux tools are preferred: The team wants to troubleshoot connectivity issues using
tcpdump,traceroute, and cloud-native flow logs.
In our case, the candid assessment was: We bought a Formula 1 race car just to pick up bread rolls in city traffic.
5. Why "Boring Infrastructure" Wins in Production
In his renowned essay "Choose Boring Technology", Dan McKinley argues that every engineering team has a finite budget of "innovation tokens." Spending those tokens on base infrastructure like the network interface leaves fewer resources for delivering core product value.
Kubernetes networking should behave like tap water: invisible, dependable, and boring.
Advantages of boring technology in production:
- Reduced cognitive overhead: No one needs training in eBPF bytecode to diagnose an everyday routing fault.
- Shorter Mean Time to Recovery (MTTR): During incidents, teams rely on established runbooks and globally documented troubleshooting practices.
- Resilient upgrade cycles: Managed Kubernetes upgrades (such as AWS EKS or Google GKE) proceed smoothly without requiring pre-flight testing of third-party kernel modules.
- Optimized resource utilization: Fewer heavy DaemonSets translate to more available CPU and RAM for user pods, as highlighted in our guide on Azure Kubernetes Cost Optimization.
6. Our Way Back: How We Uninstalled and Migrated Away From Cilium
The verdict was clear: Cilium had to be phased out. But how do you replace a CNI in an active production Kubernetes cluster without downtime?
Migration Steps at a Glance
Replacing a CNI in a live cluster resembles open-heart surgery. We executed the transition across four distinct phases:
# 1. Inventory existing Cilium resources
kubectl get ciliumnetworkpolicies -A
kubectl get ciliumclusterwidenetworkpolicies
# 2. Translate into standard Kubernetes NetworkPolicies
# (Replace proprietary L7 rules with standard L3/L4 egress/ingress)
# 3. Prepare the new node group with AWS VPC CNI and kube-proxy
# Nodes are provisioned while Cilium is still running on existing nodes
# 4. Gradually drain old Cilium nodes
kubectl cordon <node-name>
kubectl drain <node-name> --ignore-daemonsets --delete-emptydir-data
- Policy translation: All
CiliumNetworkPolicieswere reviewed. Because Layer-7 filtering was no longer required at the network level, we translated rules into standard KubernetesNetworkPolicyobjects. - Relocate egress monitoring to APM: For tracking external HTTP status codes, we instrumented central HTTP client libraries with OpenTelemetry. This captured not only status codes, but also tenant IDs and transaction metadata.
- Deploy parallel node pools: We provisioned new worker node groups configured with the official
aws-node(AWS VPC CNI) andkube-proxy. - Stepwise draining & CNI decommissioning: Workloads migrated progressively via
kubectl drainfrom old Cilium nodes to new AWS VPC CNI nodes. Once all pods were relocated, the Cilium Helm chart and leftover CRDs were cleanly uninstalled.
Results After the Rollback
Since returning to the standard network stack, we have seen measurable improvements:
- Zero CNI-related incident tickets: Over the past six months, not a single unexplained networking problem has occurred.
- Frictionless EKS upgrades: Cluster upgrades to the latest Kubernetes versions executed automatically without operational issues.
- Reclaimed resources: Each node freed approximately 200–400 MB of RAM alongside CPU cycles previously allocated to Cilium agents and BPF maps.
- Streamlined onboarding: New engineers encounter a standard EKS environment and can contribute immediately without specialized networking onboarding.
7. Conclusion: Choosing the Right Complexity for Your Team
Cilium remains a remarkable milestone in cloud-native computing, and eBPF is undeniably a defining technology for modern infrastructure. If you manage massive clusters, run telecommunication workloads, or implement BGP peering in private data centers, Cilium is a fantastic choice.
However, for the vast majority of engineering organizations running small to medium workloads on AWS, Azure, or Google Cloud: Technical minimalism is a feature, not a deficiency.
The essential question before making any architecture decision is not: "What can this tool do?", but rather: "Can we justify the operational overhead of running this tool for the next five years?"
Sometimes, the most prudent engineering decision is not adopting the latest high-end technology—it is knowing when to say goodbye.
8. Frequently Asked Questions (FAQ)
Why did we decide against Cilium in our Kubernetes cluster?
The operational complexity of eBPF, altered debugging workflows (e.g., cilium-dbg instead of familiar tcpdump), upgrade risks during node and kernel updates, and the additional resource overhead of Hubble were disproportionate to our actual requirements in an AWS EKS cluster.
What alternative to Cilium do we use now on AWS EKS?
We rely on the native AWS VPC CNI combined with standard Kubernetes components (kube-proxy). For Layer-7 observability of external API calls, we use application-level Application Performance Monitoring (APM) and OpenTelemetry instead of kernel-level tracing.
What disadvantages does eBPF introduce to daily Kubernetes operations?
eBPF hooks deeply into the Linux kernel and bypasses traditional networking paths like iptables and conntrack. Standard tools like tcpdump do not capture traffic as expected. Troubleshooting requires specialized knowledge of eBPF maps, BPF bytecode, and utilities like bpftool or cilium-dbg.
Is Cilium generally a bad CNI solution?
Not at all. Cilium is an exceptional, cutting-edge technology. For very large clusters with thousands of services, multi-cluster mesh requirements, or strict Layer-7 security policies, Cilium is often unbeatable. However, for typical small to medium cloud clusters, it is frequently overengineering.
How difficult is uninstalling or migrating away from Cilium?
Migration requires careful planning: translating existing CiliumNetworkPolicies into standard Kubernetes NetworkPolicy resources, gradually draining node pools, cleaning up eBPF mounts (/sys/fs/bpf), and installing the target CNI (e.g., AWS VPC CNI) followed by thorough end-to-end routing validation.
What does the "Choose Boring Technology" principle mean in Kubernetes?
It emphasizes selecting established, well-understood, low-maintenance tools for critical infrastructure foundations. When an incident occurs in the middle of the night, being able to resolve it quickly with familiar standard tools matters far more than the theoretical elegance of the latest tech.
Looking for guidance to optimize your Kubernetes architecture or want to avoid overengineering in your cluster? Learn more about our Cloud-Native Consulting, our services for Migration to Kubernetes, or get in touch with our experts for Docker & Kubernetes.




















