Skip to content

High memory consumption due to ConfigMap watches #967

Description

@jotak

Before anything, note than I am not a datadog user: I'm a developer of another OLM-based operator and, while investigating memory issues, out of curiosity I wanted to test a bunch of other operators to see who else was impacted by the same issue, and it seems datadog operator is. I haven't done a deep investigation on datadog-operator in particular, if you think that this is a false-positive then I apologize for the inconvenience and you can close this issue.

Describe what happened:

I ran a simple test: installing a bunch of operators, monitoring memory consumption, creating a dummy namespace with many configmaps within. On some operators, the memory consumption remained stable; on others like this one, it increased linearly with the created configmaps.

Capture d’écran du 2023-10-27 09-03-34

This could be on purpose but my assumption is that there is little chance that your operator actually needs to watch every configmaps (is it correct?). This is a quite common problem that has been documented here: https://sdk.operatorframework.io/docs/best-practices/designing-lean-operators/#overview :

"One of the pitfalls that many operators are failing into is that they watch resources with high cardinality like secrets possibly in all namespaces. This has a massive impact on the memory used by the controller on big clusters."

From my experience, with some customers it counts in gigabytes of overhead. And I would add that it's not only about memory usage, it's also stressing Kube API with a lot of traffic.

The article above suggests a remediation using cache configuration: if this would solve the problem for you, that's great!

But in case it's more complicated, you might want to chime in here: kubernetes-sigs/controller-runtime#2570 . I'm proposing to add to controller-runtime more possibilities regarding cache management, but for that I would like to probe a bit the different use cases among OLM users, in order to understand if the solution that I'm suggesting would help others or not. I guess the goal is to find a solution that suits for the most of OLM-based operators that still struggle with that, rather than each implementing its own custom cache management.

Steps to reproduce the issue:

  • Install the operator
  • Watch memory used
  • kubectl create namespace test
  • for i in {1..500}; do kubectl create cm test-cm-$i -n test --from-file=<INSERT BIG FILE HERE> ; done

Additional environment details (Operating System, Cloud provider, etc):

Using Openshift 4.14 on AWS

Metadata

Metadata

Assignees

No one assigned

    Labels

    auto-closedIssue automatically closed due to inactivitybugSomething isn't workingstale

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions