<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Stian Øvrevåge</title><description>Technical deep dives, random thoughts, photos and video from Norway.</description><link>https://blog.stian.omg.lol/</link><language>en</language><atom:link href="https://blog.stian.omg.lol/rss.xml" rel="self" type="application/rss+xml"/><item><title>Configuring Envoy as an edge proxy - through istio</title><link>https://blog.stian.omg.lol/p/configuring-envoy-as-an-edge-proxy-through-istio/</link><guid isPermaLink="true">https://blog.stian.omg.lol/p/configuring-envoy-as-an-edge-proxy-through-istio/</guid><description>How we implemented Envoy edge proxy best practices in istio-ingressgateway.</description><pubDate>Mon, 25 Nov 2024 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://blog.stian.omg.lol/_astro/envoy-logo.D1Wv7Y0p_Z12gzGR.svg&quot; width=&quot;522&quot; height=&quot;169&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;How we implemented Envoy edge proxy best practices in istio-ingressgateway.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#introduction&quot;&gt;Introduction&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#custom-bootstrap&quot;&gt;Configuring overload manager and global connection limits using a custom Envoy bootstrap&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#envoyfilter&quot;&gt;Configuring buffer sizes and connection timeouts via EnvoyFilter&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#appendices&quot;&gt;Appendices&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#appendix-get-envoy-config&quot;&gt;Appendix A - Displaying currently active Envoy edge configuration settings&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#appendix-overload-manager-metrics&quot;&gt;Appendix B - Overload manager metrics&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#appendix-istio-at-signicat&quot;&gt;Appendix C - istio installation and configuration at Signicat&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#outro&quot;&gt;Outro&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;This is a cross-post of a blog post also published on the &lt;a href=&quot;&quot;&gt;Signicat Blog&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We deployed this with Istio 1.23 and the Envoy edge proxy recommendations as of November 2024.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;If you are using istio as a Service Mesh for your Kubernetes clusters, chances are you are also using &lt;code&gt;istio-ingressgateway&lt;/code&gt; to handle incoming traffic from the Internet.&lt;/p&gt;
&lt;p&gt;Envoy however, which istio relies on, is not tuned for running at the edge by default:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Envoy is a production-ready edge proxy, however, the default settings are tailored for the service mesh use case, and some values need to be adjusted when using Envoy as an edge proxy.
&lt;br&gt;— &lt;cite&gt;&lt;a href=&quot;https://www.envoyproxy.io/docs/envoy/latest/configuration/best_practices/edge&quot;&gt;envoyproxy.io/docs&lt;/a&gt;&lt;/cite&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The &lt;a href=&quot;https://www.envoyproxy.io/docs/envoy/latest/configuration/best_practices/edge&quot;&gt;Envoy edge proxy best practices&lt;/a&gt; document outlines specific recommended configuration parameters for running envoy (and thus &lt;code&gt;istio-ingressgateway&lt;/code&gt;) on the edge.&lt;/p&gt;
&lt;p&gt;It’s not immediately obvious how you would propagate these configurations through the regular istio installation and configuration procedures. There is an &lt;a href=&quot;https://github.com/istio/istio/issues/24715&quot;&gt;open feature request on GitHub&lt;/a&gt; asking for the ability to configure &lt;code&gt;istio-ingressgateway&lt;/code&gt; according to best practices.&lt;/p&gt;
&lt;p&gt;Here I’ll show you at least one way of getting these settings deployed.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;We need two different approaches to achieve our goals. A &lt;a href=&quot;https://github.com/istio/istio/tree/master/samples/custom-bootstrap&quot;&gt;custom bootstrap configuration&lt;/a&gt; and an &lt;a href=&quot;https://istio.io/latest/docs/reference/config/networking/envoy-filter/&quot;&gt;EnvoyFilter&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;custom-bootstrap&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;configuring-overload-manager-and-global-connection-limits-using-a-custom-envoy-bootstrap&quot;&gt;Configuring overload manager and global connection limits using a custom Envoy bootstrap&lt;/h2&gt;
&lt;p&gt;To enable/configure &lt;a href=&quot;https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/operations/overload_manager&quot;&gt;Envoy overload manager&lt;/a&gt; and global connection limits we first create our config file:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;yaml&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;apiVersion&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;v1&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;kind&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;ConfigMap&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;metadata&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;  name&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;istio-envoy-custom-bootstrap-config&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;  namespace&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;istio-system&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;data&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;  custom_bootstrap.yaml&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;    # Untrusted downstreams:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;    overload_manager:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      refresh_interval: 0.25s&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      resource_monitors:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      - name: &quot;envoy.resource_monitors.fixed_heap&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;        typed_config:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          &quot;@type&quot;: type.googleapis.com/envoy.extensions.resource_monitors.fixed_heap.v3.FixedHeapConfig&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          max_heap_size_bytes: 350000000 # 350000000=350MB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      - name: &quot;envoy.resource_monitors.global_downstream_max_connections&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;        typed_config:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          &quot;@type&quot;: type.googleapis.com/envoy.extensions.resource_monitors.downstream_connections.v3.DownstreamConnectionsConfig&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          max_active_downstream_connections: 25000&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      actions:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      # Possible actions: https://www.envoyproxy.io/docs/envoy/latest/configuration/operations/overload_manager/overload_manager#overload-actions&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      - name: &quot;envoy.overload_actions.shrink_heap&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;        triggers:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;        - name: &quot;envoy.resource_monitors.fixed_heap&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          threshold:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;            value: 0.9&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      - name: &quot;envoy.overload_actions.stop_accepting_requests&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;        triggers:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;        - name: &quot;envoy.resource_monitors.fixed_heap&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          threshold:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;            value: 0.95&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      # Additional settings from https://www.envoyproxy.io/docs/envoy/latest/configuration/operations/overload_manager/overload_manager&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      - name: &quot;envoy.overload_actions.disable_http_keepalive&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;        triggers:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          - name: &quot;envoy.resource_monitors.fixed_heap&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;            threshold:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;              value: 0.95&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      # From https://www.envoyproxy.io/docs/envoy/latest/configuration/operations/overload_manager/overload_manager#reducing-timeouts&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      - name: &quot;envoy.overload_actions.reduce_timeouts&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;        triggers:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          - name: &quot;envoy.resource_monitors.fixed_heap&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;            scaled:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;              scaling_threshold: 0.85&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;              saturation_threshold: 0.95&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;        typed_config:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          &quot;@type&quot;: type.googleapis.com/envoy.config.overload.v3.ScaleTimersOverloadActionConfig&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          timer_scale_factors:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;            - timer: HTTP_DOWNSTREAM_CONNECTION_IDLE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;              min_timeout: 2s&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      # https://www.envoyproxy.io/docs/envoy/latest/configuration/operations/overload_manager/overload_manager#load-shed-points&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      loadshed_points:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;        - name: &quot;envoy.load_shed_points.tcp_listener_accept&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          triggers:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;            - name: &quot;envoy.resource_monitors.fixed_heap&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;              threshold:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;                value: 0.95&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;    # From https://www.envoyproxy.io/docs/envoy/latest/configuration/best_practices/edge#best-practices-edge / https://www.envoyproxy.io/docs/envoy/latest/configuration/listeners/runtime#config-listeners-runtime&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;    # Also mentioned in https://istio.io/latest/news/security/istio-security-2020-007/#mitigation with much higher limits&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;    # Here we configure one limit for the &quot;regular&quot; public listener on port 8443 and a separate global limit that is higher to&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;    # avoid starving connections for admin and metrics and probes&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;    layered_runtime:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      layers:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      - name: static_layer_0&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;        static_layer:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          envoy:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;            resource_limits:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;              listener:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;                0.0.0.0_8443:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;                  connection_limit: 10000&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Let’s save it to &lt;code&gt;istio-envoy-custom-bootstrap-config.yaml&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;There is a field here you MUST adjust to your environment.&lt;/strong&gt; That is the &lt;code&gt;max_heap_size_bytes&lt;/code&gt; which we set to about 90% of the configured K8s memory limit.&lt;/p&gt;
&lt;p&gt;What this does is inform the overload manager of how much memory it has available, and is used for evaluating percentage of current usage compared to what it thinks it has available, that again triggers overload actions at certain thresholds.&lt;/p&gt;
&lt;p&gt;You may also have to adjust the second to last line (&lt;code&gt;0.0.0.0_8443&lt;/code&gt;) in case your public listener is named something else.&lt;/p&gt;
&lt;p&gt;Now we install it in the cluster:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;kubectl&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; apply&lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt; -n&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; istio-system&lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt; -f&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; istio-envoy-custom-bootstrap-config.yaml&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Then we can use an &lt;a href=&quot;https://istio.io/latest/docs/setup/additional-setup/customize-installation/&quot;&gt;overlay&lt;/a&gt; to modify the Deployment that &lt;code&gt;istioctl&lt;/code&gt; produces, before &lt;code&gt;istioctl&lt;/code&gt; actually installs it in the cluster. This is the path and contents of what you need to add to your existing IstioOperator that you feed to &lt;code&gt;istioctl&lt;/code&gt;:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;yaml&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;apiVersion&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;install.istio.io/v1alpha1&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;kind&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;IstioOperator&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;spec&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;  components&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:        &lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;    ingressGateways&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;      - &lt;/span&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;name&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;istio-ingressgateway&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;        k8s&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;          overlays&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;            - &lt;/span&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;kind&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;Deployment&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;              name&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;istio-ingressgateway&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;              patches&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;                - &lt;/span&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;path&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;spec.template.spec.containers.[name:istio-proxy].env[-1]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;                  value&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;                    name&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;ISTIO_BOOTSTRAP_OVERRIDE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;                    value&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;/etc/istio/custom-bootstrap/custom_bootstrap.yaml&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;                - &lt;/span&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;path&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;spec.template.spec.containers.[name:istio-proxy].volumeMounts[-1]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;                  value&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;                    mountPath&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;/etc/istio/custom-bootstrap&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;                    name&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;custom-bootstrap-volume&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;                    readOnly&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;true&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;                - &lt;/span&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;path&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;spec.template.spec.volumes[-1]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;                  value&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;                    configMap&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;                      name&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;istio-envoy-custom-bootstrap-config&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;                      defaultMode&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;420&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;                      optional&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;false&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;                    name&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;custom-bootstrap-volume&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;Huge shoutout to our eminent &lt;a href=&quot;https://www.linkedin.com/in/csaba-k%C3%A1rp%C3%A1ti-4a229086/&quot;&gt;Csaba Kárpáti&lt;/a&gt; which helped me out actually getting the overlay above work.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;envoyfilter&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;configuring-buffer-sizes-and-connection-timeouts-via-envoyfilter&quot;&gt;Configuring buffer sizes and connection timeouts via EnvoyFilter&lt;/h2&gt;
&lt;p&gt;We set the rest of the recommended configurations via an EnvoyFilter.&lt;/p&gt;
&lt;p&gt;Create a file for it, named &lt;code&gt;listener-filters-edge.yaml&lt;/code&gt; for example:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;yaml&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;# Based on recommendations for edge deployments with untrusted downstreams:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;# - https://www.envoyproxy.io/docs/envoy/latest/configuration/best_practices/edge#best-practices-edge&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;# - https://www.envoyproxy.io/docs/envoy/latest/faq/configuration/timeouts#faq-configuration-timeouts&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;apiVersion&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;networking.istio.io/v1alpha3&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;kind&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;EnvoyFilter&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;metadata&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;  name&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;listener-filters-edge&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;spec&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;  workloadSelector&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;    labels&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;      istio&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;ingressgateway&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;  configPatches&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;    - &lt;/span&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;applyTo&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;LISTENER&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;      match&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;        context&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;GATEWAY&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;      patch&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;        operation&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;MERGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;        value&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;          per_connection_buffer_limit_bytes&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;32768&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt; # Doc examples 32 KiB # Default 1MB&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;    - &lt;/span&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;applyTo&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;NETWORK_FILTER&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;      match&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;        context&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;GATEWAY&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;        listener&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;          filterChain&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;            filter&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;              name&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;&quot;envoy.filters.network.http_connection_manager&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;      patch&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;        operation&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;MERGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;        value&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;          name&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;&quot;envoy.filters.network.http_connection_manager&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;          typed_config&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;            &quot;@type&quot;&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;&quot;type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;            # https://www.envoyproxy.io/docs/envoy/latest/api-v3/extensions/filters/network/http_connection_manager/v3/http_connection_manager.proto#envoy-v3-api-field-extensions-filters-network-http-connection-manager-v3-httpconnectionmanager-request-headers-timeout&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;            request_headers_timeout&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;10s&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;                         # Default no timeout&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;            # https://www.envoyproxy.io/docs/envoy/latest/api-v3/config/core/v3/protocol.proto#envoy-v3-api-msg-config-core-v3-http1protocoloptions&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;            common_http_protocol_options&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;              max_connection_duration&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;60s&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;                        # Default no timeout&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;              idle_timeout&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;900s&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;                                  # Default 1 hour. Doc example 900s&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;              headers_with_underscores_action&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;REJECT_REQUEST&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;            # https://www.envoyproxy.io/docs/envoy/latest/api-v3/config/core/v3/protocol.proto#config-core-v3-http2protocoloptions&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;            http2_protocol_options&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;              max_concurrent_streams&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;100&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;                         # Default 2147483647&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;              initial_stream_window_size&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;65536&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;                   # Doc examples 64 KiB - Default 268435456 (256 * 1024 * 1024)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;              initial_connection_window_size&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;1048576&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;             # Doc examples 1 MiB - Same default as initial_stream_window_size&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;            # https://www.envoyproxy.io/docs/envoy/latest/api-v3/extensions/filters/network/http_connection_manager/v3/http_connection_manager.proto.html#extensions-filters-network-http-connection-manager-v3-httpconnectionmanager&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;            stream_idle_timeout&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;300s&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;               # Default 5 mins. Must be disabled for long-lived and streaming requests&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;            request_timeout&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;300s&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6A737D;--shiki-dark:#6A737D&quot;&gt;                   # Default no timeout. Must be disabled for long-lived and streaming requests&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;            use_remote_address&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;true&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;            normalize_path&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;true&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;            merge_slashes&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;true&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;            path_with_escaped_slashes_action&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;UNESCAPE_AND_REDIRECT&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And install it like usual:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;kubectl&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; apply&lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt; -n&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; istio-system&lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt; -f&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; listener-filters-edge.yaml&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;appendices&quot;&gt;Appendices&lt;/h2&gt;
&lt;p&gt;&lt;a id=&quot;appendix-get-envoy-config&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;appendix-a---displaying-currently-active-envoy-edge-configuration-settings&quot;&gt;Appendix A - Displaying currently active Envoy edge configuration settings&lt;/h3&gt;
&lt;p&gt;Here is a handy script for printing the currently active settings to console. Useful for verifying the changes have actually made it all the way to Envoy.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;bash&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;CONFIG_FILE&lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;igw_config.json&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;echo&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;Looking for pods labeled istio=ingressgateway&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;ISTIO_INGRESSGATEWAY_POD&lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;$(&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;kubectl&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; get&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; pods&lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt; -n&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; istio-system&lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt; -l&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; istio=ingressgateway&lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt; -o&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; jsonpath=&apos;{.items[0].metadata.name}&apos;&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;echo&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;Using &lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;$ISTIO_INGRESSGATEWAY_POD&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; and dumping configuration to &lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;$CONFIG_FILE&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;kubectl&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; exec&lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt; -n&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; istio-system&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $ISTIO_INGRESSGATEWAY_POD &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;curl&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; http://localhost:15000/config_dump&lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt; &amp;gt;&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;echo&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;Custom bootstrap configuration: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;bootstrap.overload_manager.refresh_interval: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt; -r&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[0].bootstrap.overload_manager.refresh_interval&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;max_active_downstream_connections: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[0].bootstrap.overload_manager.resource_monitors[] | select(.name == &quot;envoy.resource_monitors.global_downstream_max_connections&quot;) | .typed_config.max_active_downstream_connections&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;max_heap_size_bytes: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[0].bootstrap.overload_manager.resource_monitors[] | select(.name == &quot;envoy.resource_monitors.fixed_heap&quot;) | .typed_config.max_heap_size_bytes&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;overload_actions.shrink_heap: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[0].bootstrap.overload_manager.actions[] | select(.name == &quot;envoy.overload_actions.shrink_heap&quot;) | .triggers[0].threshold.value&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;overload_actions.stop_accepting_requests: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[0].bootstrap.overload_manager.actions[] | select(.name == &quot;envoy.overload_actions.stop_accepting_requests&quot;) | .triggers[0].threshold.value&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;overload_actions.disable_http_keepalive: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[0].bootstrap.overload_manager.actions[] | select(.name == &quot;envoy.overload_actions.disable_http_keepalive&quot;) | .triggers[0].threshold.value&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;overload_actions.reduce_timeouts scaling_threshold: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[0].bootstrap.overload_manager.actions[] | select(.name == &quot;envoy.overload_actions.reduce_timeouts&quot;) | .triggers[0].scaled.scaling_threshold&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;overload_actions.reduce_timeouts saturation_threshold: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[0].bootstrap.overload_manager.actions[] | select(.name == &quot;envoy.overload_actions.reduce_timeouts&quot;) | .triggers[0].scaled.saturation_threshold&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;overload_actions.reduce_timeouts timer_scale_factors timer: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[0].bootstrap.overload_manager.actions[] | select(.name == &quot;envoy.overload_actions.reduce_timeouts&quot;) | .typed_config.timer_scale_factors[0].timer&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;overload_actions.reduce_timeouts timer_scale_factors min_timeout: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[0].bootstrap.overload_manager.actions[] | select(.name == &quot;envoy.overload_actions.reduce_timeouts&quot;) | .typed_config.timer_scale_factors[0].min_timeout&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;load_shed_points.tcp_listener_accept: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[0].bootstrap.overload_manager.loadshed_points[] | select(.name == &quot;envoy.load_shed_points.tcp_listener_accept&quot;) | .triggers[0].threshold.value&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;resource_limits.0.0.0.0_8443.connection_limit: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[0].bootstrap.layered_runtime.layers[] | select(.name == &quot;static_layer_0&quot;) | .static_layer.envoy.resource_limits.listener.&quot;0.0.0.0_8443&quot;.connection_limit&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;echo&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;EnvoyFilter configuration: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;per_connection_buffer_limit_bytes: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[2].dynamic_listeners[] | select(.name == &quot;0.0.0.0_8443&quot;) | .active_state.listener.per_connection_buffer_limit_bytes&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;http2_protocol_options.max_concurrent_streams: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[2].dynamic_listeners[] | select(.name == &quot;0.0.0.0_8443&quot;) | .active_state.listener.filter_chains[0].filters[0].typed_config.http2_protocol_options.max_concurrent_streams&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;http2_protocol_options.initial_stream_window_size: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[2].dynamic_listeners[] | select(.name == &quot;0.0.0.0_8443&quot;) | .active_state.listener.filter_chains[0].filters[0].typed_config.http2_protocol_options.initial_stream_window_size&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;http2_protocol_options.initial_connection_window_size: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[2].dynamic_listeners[] | select(.name == &quot;0.0.0.0_8443&quot;) | .active_state.listener.filter_chains[0].filters[0].typed_config.http2_protocol_options.initial_connection_window_size&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;stream_idle_timeout: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[2].dynamic_listeners[] | select(.name == &quot;0.0.0.0_8443&quot;) | .active_state.listener.filter_chains[0].filters[0].typed_config.stream_idle_timeout&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;request_timeout: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[2].dynamic_listeners[] | select(.name == &quot;0.0.0.0_8443&quot;) | .active_state.listener.filter_chains[0].filters[0].typed_config.request_timeout&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;common_http_protocol_options.idle_timeout: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[2].dynamic_listeners[] | select(.name == &quot;0.0.0.0_8443&quot;) | .active_state.listener.filter_chains[0].filters[0].typed_config.common_http_protocol_options.idle_timeout&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;common_http_protocol_options.max_connection_duration: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[2].dynamic_listeners[] | select(.name == &quot;0.0.0.0_8443&quot;) | .active_state.listener.filter_chains[0].filters[0].typed_config.common_http_protocol_options.max_connection_duration&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;printf&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &quot;request_headers_timeout: &quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cat&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; $CONFIG_FILE &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt; jq&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt; &apos;.configs[2].dynamic_listeners[] | select(.name == &quot;0.0.0.0_8443&quot;) | .active_state.listener.filter_chains[0].filters[0].typed_config.request_headers_timeout&apos;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Save it to &lt;code&gt;get-edge-config-values.sh&lt;/code&gt; and run it with for example &lt;code&gt;bash get-edge-config-values.sh&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;appendix-overload-manager-metrics&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;appendix-b---overload-manager-metrics&quot;&gt;Appendix B - Overload manager metrics&lt;/h3&gt;
&lt;p&gt;Envoy can also export metrics related to overload manager, they can be enabled by adding &lt;code&gt;overload&lt;/code&gt; to approximately this location in the &lt;code&gt;IstioOperator&lt;/code&gt; spec:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;yaml&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;apiVersion&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;install.istio.io/v1alpha1&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;kind&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;IstioOperator&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;spec&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;  components&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;    ingressGateways&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;      - &lt;/span&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;name&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;istio-ingressgateway&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;        k8s&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;          podAnnotations&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;            proxy.istio.io/config&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|-&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;              proxyStatsMatcher:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;                inclusionPrefixes:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;                - &quot;overload&quot;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;appendix-istio-at-signicat&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;appendix-c---istio-installation-and-configuration-at-signicat&quot;&gt;Appendix C - istio installation and configuration at Signicat&lt;/h3&gt;
&lt;p&gt;At Signicat we use &lt;a href=&quot;https://helm.sh/&quot;&gt;helm&lt;/a&gt; to template all resources going in to our istio installation. That includes resource types like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;IstioOperator&lt;/li&gt;
&lt;li&gt;EnvoyFilters&lt;/li&gt;
&lt;li&gt;ConfigMaps&lt;/li&gt;
&lt;li&gt;Gateways&lt;/li&gt;
&lt;li&gt;Sidecars
and so on.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We don’t use &lt;code&gt;helm&lt;/code&gt; to install anything. Only to generate a set of manifests that is then used as inputs to the appropriate tools, like feeding generated &lt;code&gt;IstioOperator&lt;/code&gt; manifests to &lt;code&gt;istioctl&lt;/code&gt; and &lt;code&gt;EnvoyFilter&lt;/code&gt; manifests to &lt;code&gt;kubectl&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;This works fairly well and allows us to have the same set of base manifests with adjustable values and feature sets per environment and using standard &lt;code&gt;helm&lt;/code&gt; that most platform engineers are already familiar with.&lt;/p&gt;
&lt;p&gt;We also have a couple of other tricks to enable us to have zero-downtime blue-green upgrades to &lt;code&gt;istio-ingressgateway&lt;/code&gt; that we may cover in a future post.&lt;/p&gt;
&lt;h2 id=&quot;outro&quot;&gt;Outro&lt;/h2&gt;
&lt;p&gt;I hope this was helpful on your journey towards scale, resilience and reliability.&lt;/p&gt;
&lt;p&gt;The next chapter in this saga would be once we complete extensive load testing with the new configurations compared to the defaults as well as trying to find optimal values. And associated Grafana dashboards are always nice!&lt;/p&gt;
</content:encoded><category>tech</category><category>kubernetes</category><category>istio</category></item><item><title>Introduction</title><link>https://blog.stian.omg.lol/p/introduction/</link><guid isPermaLink="true">https://blog.stian.omg.lol/p/introduction/</guid><description>Home of technical deep dives, random thoughts and ideas of varying quality.</description><pubDate>Sun, 24 Nov 2024 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://blog.stian.omg.lol/_astro/20241107-_DSF2649.ByIfV5El_1wzdtC.jpg&quot; width=&quot;1600&quot; height=&quot;900&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;Home of technical deep dives, random thoughts and ideas of varying quality.&lt;/p&gt;
&lt;h4 id=&quot;my-name-is-stian-øvrevåge-and-im-a-passionate-technologist-and-wannabe-adventurer&quot;&gt;My name is Stian Øvrevåge and I’m a passionate technologist and wannabe adventurer.&lt;/h4&gt;
&lt;p&gt;In addition to posting on average one blog post a year I do other stuff related to aviation, photography and general nerding out.&lt;/p&gt;
&lt;h4 id=&quot;the-canonical-overview-of-my-online-presence-is-located-at-httpsstianomglol&quot;&gt;The canonical overview of my online presence is located at &lt;a href=&quot;https://stian.omg.lol/&quot;&gt;https://stian.omg.lol/&lt;/a&gt;.&lt;/h4&gt;
</content:encoded><category>personal</category><category>meta</category></item><item><title>Summer of 2024</title><link>https://blog.stian.omg.lol/p/summer-of-2024/</link><guid isPermaLink="true">https://blog.stian.omg.lol/p/summer-of-2024/</guid><description>Well, as you can see, mountains and museums and stuff.</description><pubDate>Sun, 21 Jul 2024 18:51:01 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/f6f369ef52209d3d/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/43ff51f53e12a382/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/33268b77c5d47dd2/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/5b4f68fc4bde6a86/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/0f8bf195bd983e74/w3200.jpg&quot; width=&quot;6240&quot; height=&quot;3814&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/ae662393903f56eb/w3200.jpg&quot; width=&quot;6240&quot; height=&quot;3622&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/5b027cd848f84724/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/499df639e611278b/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/20142cb8742e557d/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/90bfac7548054bdc/w3200.jpg&quot; width=&quot;4057&quot; height=&quot;6086&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/e2b14aaacfdb8ac5/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/df0909aa18f705a3/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/7f591ef77c81434f/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/9c714a2a79fe62e7/w3200.jpg&quot; width=&quot;6240&quot; height=&quot;4160&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/4550b8c760f24ffc/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/ab9f0b5fa2d9cb8a/w3200.jpg&quot; width=&quot;3927&quot; height=&quot;5890&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/d52a1a81be007a12/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/5a26e44c43976b4d/w3200.jpg&quot; width=&quot;6240&quot; height=&quot;4160&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/acdb8d18a10e9b1f/w3200.jpg&quot; width=&quot;6240&quot; height=&quot;4160&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/e984cade9bfb6c6d/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/b8e9dca4991246d1/w3200.jpg&quot; width=&quot;6240&quot; height=&quot;4160&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/ae041c3ae43ed2a0/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/edb9ae8b84564898/w3200.jpg&quot; width=&quot;6240&quot; height=&quot;4160&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/dd0463a71749bee5/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/6d0a747e077650cc/w3200.jpg&quot; width=&quot;6240&quot; height=&quot;4160&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/266d6c5014a7d76a/w3200.jpg&quot; width=&quot;6240&quot; height=&quot;4160&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/c6ff01e2c918e709/w3200.jpg&quot; width=&quot;6240&quot; height=&quot;4160&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/d06bcf26ff3f8734/w3200.jpg&quot; width=&quot;6240&quot; height=&quot;4160&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/771950b45d2e88c9/w3200.jpg&quot; width=&quot;6240&quot; height=&quot;4160&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/bc561e63c3789720/w3200.jpg&quot; width=&quot;4160&quot; height=&quot;6240&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;Well, as you can see, mountains and museums and stuff.&lt;/p&gt;
</content:encoded><category>art</category><category>photography</category></item><item><title>Snapshots of 2024</title><link>https://blog.stian.omg.lol/p/snapshots-of-2024/</link><guid isPermaLink="true">https://blog.stian.omg.lol/p/snapshots-of-2024/</guid><description>More random photos from 2024.</description><pubDate>Wed, 29 May 2024 15:17:03 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/d45e9284de97a287/w1080.jpg&quot; width=&quot;1080&quot; height=&quot;1350&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/8f3483bcb708eed8/w1080.jpg&quot; width=&quot;1080&quot; height=&quot;1350&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/0f8d7971cb6176ed/w1080.jpg&quot; width=&quot;1080&quot; height=&quot;1350&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/ac41ab1dcb6b1f18/w1080.jpg&quot; width=&quot;1080&quot; height=&quot;1350&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/d611af1b06ec1fd3/w1080.jpg&quot; width=&quot;1080&quot; height=&quot;1350&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/f9355b9d441a5218/w1080.jpg&quot; width=&quot;1080&quot; height=&quot;1350&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/fabd8475cee80514/w1080.jpg&quot; width=&quot;1080&quot; height=&quot;1350&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/82829fd62d689191/w1080.jpg&quot; width=&quot;1080&quot; height=&quot;1350&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://media.skypony.no/d/bc45e9b05b4d5ddb/w1080.jpg&quot; width=&quot;1080&quot; height=&quot;1350&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;More random photos from 2024.&lt;/p&gt;
</content:encoded><category>art</category><category>photography</category></item><item><title>Kubernetes: Don&apos;t use NodeLocal DNSCache on GKE (without Cloud DNS)</title><link>https://blog.stian.omg.lol/p/kubernetes-dont-use-nodelocal-dnscache-on-gke-without-cloud-dns/</link><guid isPermaLink="true">https://blog.stian.omg.lol/p/kubernetes-dont-use-nodelocal-dnscache-on-gke-without-cloud-dns/</guid><description>A travel diary of how we diagnosed a latency spike problem on GKE and the surprising things we learned along the way.</description><pubDate>Wed, 13 Dec 2023 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://blog.stian.omg.lol/_astro/cover.C10vK77J_Z2mjLNH.jpg&quot; width=&quot;1024&quot; height=&quot;585&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;A travel diary of how we diagnosed a latency spike problem on GKE and the surprising things we learned along the way.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#introduction&quot;&gt;Introduction&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#tl-dr&quot;&gt;TL;DR:&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#nodelocal-dnscache-coredns&quot;&gt;NodeLocal DNSCache / CoreDNS&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#kubedns-dnsmasq&quot;&gt;kube-dns / dnsmasq&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#am-i-affected&quot;&gt;Am I affected?&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#investigation&quot;&gt;Investigation&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#the-initial-problem&quot;&gt;The initial Problem&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#istio-envoy-metrics&quot;&gt;Istio / Envoy metrics&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#reproducing&quot;&gt;Reproducing and troubleshooting HTTP connection problems&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#troubleshooting-dns&quot;&gt;Troubleshooting potentially slow or broken DNS lookups&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#appendices&quot;&gt;Appendices&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#appendix-enabling-additional-envoy-metrics&quot;&gt;Appendix - Enabling additional envoy metrics&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#appendix-overview-of-dns-on-gke&quot;&gt;Appendix - Overview of DNS on GKE&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#appendix-reducing-dns-lookups&quot;&gt;Appendix - Reducing DNS lookups in Kubernetes and GKE&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#appendix-overview-of-gke-dns-with-nodelocal-dnscache&quot;&gt;Appendix - Overview of GKE DNS with NodeLocal DNSCache&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#appendix-enabling-node-local-dns-metrics&quot;&gt;Appendix - Enabling node-local-dns metrics&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#appendix-using-dnsperf&quot;&gt;Appendix - Using dnsperf to test DNS performance&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#appendix-network-packet-capture&quot;&gt;Appendix - Network packet capture without dependencies&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#appendix-verbose-logging-on-kube-dns&quot;&gt;Appendix - Verbose logging on kube-dns&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#appendix-analyzing-dnsmasq-logs&quot;&gt;Appendix - Analyzing dnsmasq logs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#appendix-analyzing-concurrent-tcp-connections&quot;&gt;Appendix - Analyzing concurrent TCP connections&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#analyzing-dns-problems-based-on-packet-capture&quot;&gt;Appendix - Analyzing DNS problems based on packet captures&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#footnotes&quot;&gt;Footnotes&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;This is a cross-post of a blog post also published on the &lt;a href=&quot;https://www.signicat.com/blog/dont-use-nodelocal-dnscache-on-gke&quot;&gt;Signicat Blog&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Last year I was freelancing as a consultant in Signicat and recently I returned, now as an actual employee!&lt;/p&gt;
&lt;p&gt;The first week after returning, a tech lead for another team reaches out to me about a technical issue he’s been struggling with.&lt;/p&gt;
&lt;p&gt;In this blog post I’ll try to guide you through the troubleshooting with actionable take-aways. It appears it’s going to be a long one with a lot of detours, so I’ve summarized our findings and recommendations here on the top starting right below this introduction. There are a few appendixes at the end that I’ve tried to make self-contained as to make them useful in other contexts and without necessarily having to follow the main story.&lt;/p&gt;
&lt;p&gt;If you don’t want any spoilers but follow the bumpy journey from start to end, fill your coffee mug and skip ahead to &lt;a href=&quot;#investigation&quot;&gt;Investigation&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;tl-dr&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;tldr&quot;&gt;TL;DR:&lt;/h2&gt;
&lt;p&gt;Due to an unfortunate combination of behaviour of &lt;a href=&quot;https://coredns.io/manual/toc/&quot;&gt;CoreDNS&lt;/a&gt; (which NodeLocal DNSCache uses) and &lt;a href=&quot;https://github.com/kubernetes/dns&quot;&gt;kube-dns&lt;/a&gt; (which is the default on GKE) &lt;strong&gt;I recommend NOT using them in combination.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Since &lt;a href=&quot;https://cloud.google.com/knowledge/kb/how-to-run-coredns-on-kubernetes-engine-000004698&quot;&gt;GKE does not offer CoreDNS as a managed option&lt;/a&gt; for &lt;code&gt;kube-dns&lt;/code&gt; (even though &lt;a href=&quot;https://kubernetes.io/blog/2018/12/03/kubernetes-1-13-release-announcement/&quot;&gt;Kubernetes made CoreDNS the default in 1.13 in 2018&lt;/a&gt;) you are left with two options:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Not enabling &lt;a href=&quot;https://cloud.google.com/kubernetes-engine/docs/how-to/nodelocal-dns-cache&quot;&gt;NodeLocal DNSCache on GKE&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Switching from &lt;code&gt;kube-dns&lt;/code&gt; &lt;a href=&quot;https://cloud.google.com/kubernetes-engine/docs/how-to/cloud-dns&quot;&gt;to Google Cloud DNS&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;What is this unfortunate combination you ask?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;nodelocal-dnscache-coredns&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;nodelocal-dnscache--coredns&quot;&gt;NodeLocal DNSCache / CoreDNS&lt;/h3&gt;
&lt;p&gt;NodeLocal DNSCache (in GKE at least) is configured with:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# kubectl get configmap --namespace kube-system node-local-dns -o yaml | yq .data.Corefile&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cluster.local:53 {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  forward . __PILLAR__CLUSTER__DNS__ {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      force_tcp&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      expire 1s&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This means CoreDNS will upgrade any and all incoming DNS requests to TCP connections before connecting to &lt;code&gt;dnsmasq&lt;/code&gt; (the first of two containers in a &lt;code&gt;kube-dns&lt;/code&gt; Pod). CoreDNS reuses TCP connections if available. Unused connections in the connection pool should be cleaned up every 1 second (in theory). If no connections are available a new will be created, apparently with no upper bound.&lt;/p&gt;
&lt;p&gt;This means new connections may be created en-masse when needed, but old connections can take a while (I’ve observed 5-10 seconds) before being cleaned up, and most importantly, closed.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;kubedns-dnsmasq&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;kube-dns--dnsmasq&quot;&gt;kube-dns / dnsmasq&lt;/h3&gt;
&lt;p&gt;&lt;code&gt;dnsmasq&lt;/code&gt; has a &lt;a href=&quot;https://github.com/imp/dnsmasq/blob/master/src/config.h#L18&quot;&gt;hardcoded maximum number of 20 workers&lt;/a&gt;. For us that means each &lt;code&gt;kube-dns&lt;/code&gt; Pod is limited to 20 open connections.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;kube-dns&lt;/code&gt; scaling is managed by a bespoke kube-dns autoscaler that by default in GKE is &lt;a href=&quot;https://cloud.google.com/kubernetes-engine/docs/how-to/nodelocal-dns-cache#scaling_up_kube-dns&quot;&gt;configured&lt;/a&gt; like:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# kubectl get configmap --namespace kube-system kube-dns-autoscaler -o yaml | yq .data.linear&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;{&quot;coresPerReplica&quot;:256, &quot;nodesPerReplica&quot;:16,&quot;preventSinglePointFailure&quot;:true}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;preventSinglePointFailure&lt;/code&gt; - Run at least two Pods.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;nodesPerReplica&lt;/code&gt; - Run one &lt;code&gt;kube-dns&lt;/code&gt; Pod for each 16 nodes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This default setup will have two &lt;code&gt;kube-dns&lt;/code&gt; Pods until your cluster grows beyond 32 nodes.&lt;/p&gt;
&lt;p&gt;Two &lt;code&gt;kube-dns&lt;/code&gt; Pods have a limit of 40 open TCP connections in total from the &lt;code&gt;node-local-dns&lt;/code&gt; Pods running on each node. The &lt;code&gt;node-local-dns&lt;/code&gt; Pods though are happy to try to open many more TCP connections.&lt;/p&gt;
&lt;p&gt;GKE kube-dns docs mention “&lt;a href=&quot;https://cloud.google.com/kubernetes-engine/docs/how-to/kube-dns#performance_limitations_with_kube-dns&quot;&gt;Performance limitations with kube-dns&lt;/a&gt;” as a known issue suggesting enabling NodeLocal DNSCache as a potential fix.&lt;/p&gt;
&lt;p&gt;And GKE NodeLocal DNSCache docs mention “&lt;a href=&quot;https://cloud.google.com/kubernetes-engine/docs/how-to/nodelocal-dns-cache#timeout_issues&quot;&gt;NodeLocal DNSCache timeout errors&lt;/a&gt;” as a known issue with the possible reasons being&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;An underlying network connectivity problem. (&lt;em&gt;Spoiler: it probably isn’t&lt;/em&gt;)&lt;/li&gt;
&lt;li&gt;Significantly increased DNS queries from the workload or due to node pool upscaling. (&lt;em&gt;Spoiler: it probably isn’t&lt;/em&gt;)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;With the fix being&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“The workaround is to increase the number of kube-dns replicas by tuning the autoscaling parameters.”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Increasing the number of &lt;code&gt;kube-dns&lt;/code&gt; Pods will reduce the frequency and impact but not eliminate it. Until we migrate to Cloud DNS we run &lt;code&gt;kube-dns&lt;/code&gt; at a ratio of 1:1.5 of nodes by configuring autoscaling with &lt;code&gt;&quot;coresPerReplica&quot;:24&lt;/code&gt; and using 16 core nodes. Resulting in 12 &lt;code&gt;kube-dns&lt;/code&gt; Pods in a 17 node cluster.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;By looking at the &lt;code&gt;coredns_forward_conn_cache_misses_total&lt;/code&gt; metric I observe it at least increasing by more than 300 in a 15 second metric sampling window for &lt;em&gt;one&lt;/em&gt; &lt;code&gt;node-local-dns&lt;/code&gt; Pod. This means &lt;em&gt;on average&lt;/em&gt; during those 15 seconds 20 new TCP connections were attempted since no existing connection could be re-used. (&lt;a href=&quot;#footnote-a&quot;&gt;Footnote A&lt;/a&gt;)&lt;/p&gt;
&lt;p&gt;This means even setting &lt;code&gt;&quot;nodesPerReplica&quot;:1&lt;/code&gt; thus running one &lt;code&gt;kube-dns&lt;/code&gt; Pod for each node may not be enough to guarantee not hitting the 20 process limit in &lt;code&gt;dnsmasq&lt;/code&gt; occasionally.&lt;/p&gt;
&lt;p&gt;I guess you could start lowering &lt;code&gt;coresPerReplica&lt;/code&gt; to have more than one &lt;code&gt;kube-dns&lt;/code&gt; Pod for each node, but now it’s getting silly.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;To re-iterate; If you’re using NodeLocal DNSCache and &lt;code&gt;kube-dns&lt;/code&gt; you should plan to migrate to Cloud DNS. You can alleviate the problem in the short term by scaling up &lt;code&gt;kube-dns&lt;/code&gt; aggressively but it will not eliminate occasional latency spikes.&lt;/p&gt;
&lt;h3 id=&quot;am-i-affected&quot;&gt;Am I affected?&lt;/h3&gt;
&lt;p&gt;How do I know if I’m affected by this?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;You are using NodeLocal DNSCache (CoreDNS) and &lt;code&gt;kube-dns&lt;/code&gt; (default on GKE)&lt;/li&gt;
&lt;li&gt;You see log lines with &lt;code&gt;i/o timeout&lt;/code&gt; in node-local-dns pods (&lt;code&gt;kubectl logs -n kube-system node-local-dns-xxxxx&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;You are hitting dnsmasq max procs. Create a graph &lt;code&gt;container_processes{namespace=&quot;kube-system&quot;, container=&quot;dnsmasq&quot;}&lt;/code&gt; or use my &lt;a href=&quot;https://github.com/signicat/blog-attachements/blob/main/2023-gke-node-local-dns-cache/files/grafana-dashboard-kube-dns.json&quot;&gt;kube-dns Grafana dashboard&lt;/a&gt;. (&lt;a href=&quot;#footnote-b&quot;&gt;Footnote B&lt;/a&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;investigation&quot;&gt;Investigation&lt;/h2&gt;
&lt;h3 id=&quot;the-initial-problem&quot;&gt;The initial Problem&lt;/h3&gt;
&lt;p&gt;We have a microservice architecture and use &lt;a href=&quot;https://www.openpolicyagent.org/&quot;&gt;Open Policy Agent (OPA)&lt;/a&gt; deployed as &lt;a href=&quot;https://kubernetes.io/blog/2015/06/the-distributed-system-toolkit-patterns/#example-1-sidecar-containers&quot;&gt;sidecars&lt;/a&gt; to process and validate request tokens.&lt;/p&gt;
&lt;p&gt;However, some services were getting timeouts and he suspected it was an issue with &lt;a href=&quot;https://istio.io/&quot;&gt;istio&lt;/a&gt;, a service mesh that provides security and observability for traffic inside Kubernetes clusters.&lt;/p&gt;
&lt;p&gt;The actual error message showing up in the logs was a generic &lt;code&gt;context deadline exceeded (Client.Timeout exceeded while awaiting headers)&lt;/code&gt; which I recognize as a Golang error probably stemming from a &lt;a href=&quot;https://pkg.go.dev/net/http&quot;&gt;http.Get()&lt;/a&gt; call or similar.&lt;/p&gt;
&lt;p&gt;The first thing that comes to mind is that last year we made a change in our istio practices. Inside our clusters we call services in other namespaces (crossing team boundaries) using the same FQDN domain and path that our customers would use, not internal Kubernetes names. So internally we’re calling &lt;code&gt;api.signicat.com/service&lt;/code&gt; and not &lt;code&gt;service.some-team-namespace&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;He had noticed that the problem seemed to have increased when doing the switch from internal DNS names to FQDN.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;istio-envoy-metrics&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;istio--envoy-metrics&quot;&gt;Istio / Envoy metrics&lt;/h3&gt;
&lt;p&gt;We started troubleshooting by enabling additional metrics in envoy to hopefully be able to confirm that the problem was visible as envoy timeouts or some other kind of error.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;a href=&quot;https://www.envoyproxy.io/&quot;&gt;Envoy&lt;/a&gt; is the actual software that traffic goes through in an istio service mesh.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This turned out to be a dead end and we couldn’t find any signs of timeouts or errors in the envoy metrics.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Enabling these additional metrics was a bit of a chore and there were a few surprises. Have a look at the &lt;a href=&quot;#appendix-enabling-additional-envoy-metrics&quot;&gt;Enabling additional envoy metrics&lt;/a&gt; appendix for steps and caveats.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;reproducing&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;reproducing-and-troubleshooting-http-connection-problems&quot;&gt;Reproducing and troubleshooting HTTP connection problems&lt;/h3&gt;
&lt;p&gt;At this point I’m thinking the problem may be something else. Failure to acquire a HTTP connection from the pool or a TCP socket maybe?&lt;/p&gt;
&lt;p&gt;Using the &lt;a href=&quot;https://pkg.go.dev/net/http/httptrace&quot;&gt;httptrace&lt;/a&gt; golang library I created a small program that would continuously query both the internal &lt;code&gt;some-service.some-namespace&lt;/code&gt; hostname and FQDN &lt;code&gt;api.signicat.com/some-service&lt;/code&gt;, while logging and counting each phase of the HTTP request as well as it’s timings.&lt;/p&gt;
&lt;p&gt;Here is a quick and dirty way of starting the program in an ephemeral Pod:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Start a Pod named `conn-test` using the `golang:latest` image. Attach to it (`-i --tty`) and start `bash`. Delete it after disconnecting (`--rm`):&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl run conn-test --rm -i --tty --image=golang:latest -- bash&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Create and enter a directory for the program:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;mkdir conn-test &amp;amp;&amp;amp; cd conn-test&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Initialize a go environment for it:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;go mod init conn-test&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Download the source:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;wget https://raw.githubusercontent.com/signicat/blog-attachements/main/2023-gke-node-local-dns-cache/files/go-http-connection-test/main.go&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Download dependencies:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;go mod tidy&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Configure it (refer to the source for more configuration options):&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;export URL_1=http://some-service.some-namespace.svc.cluster.local&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Start it:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;go run main.go&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;If you need to run it for longer it probably makes sense to build it as a Docker image, upload it to an image registry and create a Deployment for it. It also exposes metrics in Prometheus format so you can also add scraping of the metrics either through annotations or a ServiceMonitor.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;After a few hours the program crashed, and the last log message indicated that &lt;code&gt;DNSStart&lt;/code&gt; was the last thing to happened. (The program crashed since I just &lt;code&gt;log.Fatal&lt;/code&gt;ed if the request failed, which it did timing out. I since improved that, even though we already learned what we needed).&lt;/p&gt;
&lt;p&gt;We did some manual DNS lookups using &lt;code&gt;nslookup -debug api.signicat.com&lt;/code&gt; which (obviously, in retrospect) shows 6 failing DNS requests before finally getting it right:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;api.signicat.com.$namespace.svc.cluster.local - NXDOMAIN&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;api.signicat.com.svc.cluster.local - NXDOMAIN&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;api.signicat.com.cluster.local - NXDOMAIN&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;api.signicat.com.$gcp-region.c.$gcp-project.internal - NXDOMAIN&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;api.signicat.com.c.$gcp-project.internal - NXDOMAIN&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;api.signicat.com.google.internal - NXDOMAIN&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;api.signicat.com - OK&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;See &lt;a href=&quot;#appendix-overview-of-dns-on-gke&quot;&gt;Overview of DNS on GKE&lt;/a&gt; for a lengthier explanation of why this happens.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The default timeout for HTTP requests in Golang is 5 seconds. DNS resolution &lt;em&gt;should&lt;/em&gt; be fast enough and it shouldn’t be a problem to do 7 DNS lookups in that time. But if there is some sluggishness and variability in DNS resolution times the probability of having one slow lookup out of the 7 being slow increases the risk of causing a timeout. In addition, always doing 7 lookups puts additional strain on the DNS infrastructure potentially further exacerbating the probability of slow lookups.&lt;/p&gt;
&lt;p&gt;From here on we have two courses of action. Figure out if we can reduce the number of DNS lookups as well as figure out if, and why, DNS resolution isn’t consistently fast.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;See &lt;a href=&quot;#appendix-reducing-dns-lookups&quot;&gt;Reducing DNS lookups in Kubernetes and GKE&lt;/a&gt; for some thoughts on reducing these extraneous DNS lookups.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;troubleshooting-dns&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;troubleshooting-potentially-slow-or-broken-dns-lookups&quot;&gt;Troubleshooting potentially slow or broken DNS lookups&lt;/h3&gt;
&lt;p&gt;We use &lt;a href=&quot;https://kubernetes.io/docs/tasks/administer-cluster/nodelocaldns/&quot;&gt;NodeLocal DNSCache&lt;/a&gt; which on GKE is &lt;a href=&quot;https://cloud.google.com/kubernetes-engine/docs/how-to/nodelocal-dns-cache&quot;&gt;enabled by clicking the right button&lt;/a&gt;.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;NodeLocal DNSCache summary&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;It is deployed by a ReplicaSet causing one CoreDNS Pod (named &lt;code&gt;node-local-dns-x&lt;/code&gt; in &lt;code&gt;kube-system&lt;/code&gt; namespace) to be running on each node.&lt;/p&gt;
&lt;p&gt;It adds a listener on the same IP as the &lt;code&gt;kube-dns&lt;/code&gt; Service. Thereby intercepting traffic that would otherwise go to that IP.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;node-local-dns&lt;/code&gt; Pod caches both positive and negative results. Domains ending in &lt;code&gt;.cluster.local&lt;/code&gt; are forwarded to &lt;code&gt;kube-dns&lt;/code&gt; in the cluster (but through a new Service called &lt;code&gt;kube-dns-upstream&lt;/code&gt; with a different IP). Requests for other domains are forwarded to 8.8.8.8 and 8.8.4.4.&lt;/p&gt;
&lt;p&gt;NodeLocal DNSCache instances use a 2 second timeout when querying upstream DNS servers.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;See &lt;a href=&quot;#appendix-overview-of-gke-dns-with-nodelocal-dnscache&quot;&gt;Overview of GKE DNS with NodeLocal DNSCache&lt;/a&gt; for additional details.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;We wanted to look at some metrics from the &lt;code&gt;node-local-dns&lt;/code&gt; Pods but found that we didn’t have any! We fixed that but didn’t at the time learn anything new. &lt;em&gt;See &lt;a href=&quot;#appendix-enabling-node-local-dns-metrics&quot;&gt;Enabling node-local-dns metrics&lt;/a&gt; for how to fix/enable node-local-dns metrics.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Stumbling over the logs of one of the &lt;code&gt;node-local-dns-x&lt;/code&gt; Pods I notice:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;[ERROR] plugin/errors: 2 archive.ubuntu.com.some-namespace.svc.cluster.local. A: dial tcp 172.18.165.130:53: i/o timeout&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This tells us that at least lookups going to &lt;code&gt;kube-dns&lt;/code&gt; in the cluster (&lt;code&gt;172.18.165.130:53&lt;/code&gt;) are having problems.&lt;/p&gt;
&lt;p&gt;So timeouts are happening on individual lookups, it’s not just the total duration of 7 lookups timing out. And lookups are not only happening inside the client application but in &lt;code&gt;node-local-dns&lt;/code&gt; as well. We still don’t know if these lookups are lost or merely slow. But seen from the application anything longer than 2 seconds, that is timing out on &lt;code&gt;node-local-dns&lt;/code&gt;, might as well be lost.&lt;/p&gt;
&lt;p&gt;Just to be sure I checked that packets were not being dropped or lost on the nodes on both sides. And indeed the network seems fine. It’s been a long time since I’ve actually experienced problems due to packets being lost, but it’s always good to rule it out.&lt;/p&gt;
&lt;p&gt;Next step is to capture the packets as they (hopefully) leave the &lt;code&gt;node-local-dns&lt;/code&gt; Pod and (maybe) arriving at one of the &lt;code&gt;kube-dns&lt;/code&gt; Pods.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;See &lt;a href=&quot;#appendix-network-packet-capture&quot;&gt;Network packet capture without dependencies&lt;/a&gt; for how to capture packets in Kubernetes.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The packets are indeed arriving at the &lt;code&gt;dnsmasq&lt;/code&gt; container in the &lt;code&gt;kube-dns&lt;/code&gt; Pods:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Query 0x31e8:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Leaves node-local-dns eth0:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;10:53:11.134	10.0.0.107	172.20.13.4	DNS	0x31e8	Standard query 0x31e8 A archive.ubuntu.com.some-ns.svc.cluster.local&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Arrives at dnsmasq eth0:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;10:53:11.134	10.0.0.107	172.20.13.4	DNS	0x31e8	Standard query 0x31e8 A archive.ubuntu.com.some-ns.svc.cluster.local&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Looking at DNS resolution time from the &lt;code&gt;kube-dns&lt;/code&gt; container which &lt;code&gt;dnsmasq&lt;/code&gt; forwards to shows that &lt;code&gt;kube-dns&lt;/code&gt; is consistently answering queries extremely fast (not shown here).&lt;/p&gt;
&lt;p&gt;But after some head scratching I look at the time between queries arriving at &lt;code&gt;dnsmasq&lt;/code&gt; on &lt;code&gt;eth0&lt;/code&gt; before leaving again on &lt;code&gt;lo&lt;/code&gt; for &lt;code&gt;kube-dns&lt;/code&gt; and indeed there is a (relatively) long delay of 952ms between 10:53:11.134 and 10:53:12.086:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Leaves dnsmasq lo:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;10:53:12.086	127.0.0.1	127.0.0.1	DNS	0x31e8	Standard query 0x31e8 A archive.ubuntu.com.some-ns.svc.cluster.local&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;In this case the query just sits around in &lt;code&gt;dnsmasq&lt;/code&gt; for almost one second before being forwarded to &lt;code&gt;kube-dns&lt;/code&gt;!&lt;/p&gt;
&lt;p&gt;Why? As with most issues we start by checking if there are any issues with memory or CPU usage or CPU throttling (&lt;a href=&quot;#footnote-c&quot;&gt;Footnote C&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Graphs showing DNS response time on node-local-dns Pods and resource usage on kube-dns Pods.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./kube-dns-resource-usage.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;Nope. CPU and memory usage is both low and stable but the 99 percentile DNS request duration is all over the place.&lt;/p&gt;
&lt;p&gt;We also check and see that there is plenty of unused CPU available on the underlying nodes these pods are running on.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We also used &lt;code&gt;dnsperf&lt;/code&gt; to benchmark and stress-test the various components and while fun, didn’t teach us anything new. &lt;em&gt;See &lt;a href=&quot;#appendix-using-dnsperf&quot;&gt;Using dnsperf to test DNS performance&lt;/a&gt; for more information on that particular side-quest.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Next up I want to increase the logging verbosity of &lt;code&gt;dnsmasq&lt;/code&gt; to see if there are any clues to why it was (apparently) delaying processing DNS requests.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Have a look at &lt;a href=&quot;#appendix-verbose-logging-on-kube-dns&quot;&gt;Verbose logging on kube-dns&lt;/a&gt; for how to increase logging and &lt;a href=&quot;#appendix-analyzing-dnsmasq-logs&quot;&gt;Analyzing dnsmasq logs&lt;/a&gt; for how I analyzed the logs.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Analyzing the logs we learn that at least it appears that individual requests are fast and not clogging up the machinery.&lt;/p&gt;
&lt;p&gt;In the meantime Google Cloud Support came back to us asking if we could run a command &lt;code&gt;for i in $(seq 1 1800) ; do echo &quot;$(date) Try: ${i} DnsmasqProcess: $(pidof dnsmasq | wc -w)&quot;; sleep 1; done&lt;/code&gt; on the VM of the Pod. It also works to run this in a debug container attached to the &lt;code&gt;dnsmasq&lt;/code&gt; container. The output looks like:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Wed Oct 11 14:50:57 UTC 2023 Try: 1 DnsmasqProcess: 9&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Wed Oct 11 14:50:58 UTC 2023 Try: 2 DnsmasqProcess: 14&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Wed Oct 11 14:50:59 UTC 2023 Try: 3 DnsmasqProcess: 15&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Wed Oct 11 14:51:00 UTC 2023 Try: 4 DnsmasqProcess: 21&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Wed Oct 11 14:51:01 UTC 2023 Try: 5 DnsmasqProcess: 21&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;It turns out that &lt;code&gt;dnsmasq&lt;/code&gt; has a &lt;a href=&quot;https://github.com/imp/dnsmasq/blob/master/src/config.h#L18&quot;&gt;hard coded limit of 20 child processes&lt;/a&gt;. So every time we see 21 it means it will not spawn a new child to process incoming connections.&lt;/p&gt;
&lt;p&gt;I plotted this on the same graph as the DNS request times observed from raw packet captures (See &lt;a href=&quot;#analyzing-dns-problems-based-on-packet-capture&quot;&gt;Analyzing DNS problems based on packet captures&lt;/a&gt;):&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Graph showing DNS resolution time from packet captures in blue. Number of dnsmasq processes in orange.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./dnsmasq-proc-and-dns-packet-latency.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;Graphing the number of &lt;code&gt;dnsmasq&lt;/code&gt; processes and observed DNS request latency finally shows a stable correlation between our symptoms and a potential cause.&lt;/p&gt;
&lt;p&gt;You can also get a crude estimation by graphing the &lt;code&gt;container_processes{namespace=&quot;kube-system&quot;, container=&quot;dnsmasq&quot;}&lt;/code&gt; metric, or use my &lt;a href=&quot;https://github.com/signicat/blog-attachements/blob/main/2023-gke-node-local-dns-cache/files/grafana-dashboard-kube-dns.json&quot;&gt;kube-dns Grafana dashboard&lt;/a&gt; if you have cAdvisor/container metrics enabled:&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Graph showing DNS resolution time from node-local-dns and dnsmasq processes.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./dnsmasq-proc-graph-and-node-local-dns-duration.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;Going back to our packet captures I see connections that are unused for many seconds before finally being closed by &lt;code&gt;node-local-dns&lt;/code&gt;.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;08:06:07.849	172.20.9.11	10.0.0.29	53 → 46123 [ACK] Seq=1 Ack=80 Win=43648 Len=0&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;08:06:09.849	10.0.0.29	172.20.9.11	46123 → 53 [FIN, ACK] Seq=80 Ack=1 Win=42624 Len=0&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;08:06:09.890	172.20.9.11	10.0.0.29	53 → 46123 [ACK] Seq=1 Ack=81 Win=43648 Len=0&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;However &lt;code&gt;dnsmasq&lt;/code&gt; only acknowledges the request to close the connection, it does not actually close it yet. In order for the connection to be properly closed both sides have to &lt;code&gt;FIN, ACK&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;Many seconds later &lt;code&gt;dnsmasq&lt;/code&gt; faithfully tries to return a response and finish the closing of the connection, but &lt;code&gt;node-local-dns&lt;/code&gt; (CoreDNS) has promptly forgotten about the whole thing and replies with the TCP equivalent of a shrug (&lt;code&gt;RST&lt;/code&gt;):&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;08:06:15.836	172.20.9.11	10.0.0.29	53	46123	Standard query response 0x7f12 No such name A storage.googleapis.com.some-ns.svc.cluster.local SOA ns.dns.cluster.local&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;08:06:15.836	172.20.9.11	10.0.0.29	53	46123	53 → 46123 [FIN, ACK] Seq=162 Ack=81 Win=43648 Len=0&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;08:06:15.836	10.0.0.29	172.20.9.11	46123	53	46123 → 53 [RST] Seq=81 Win=0 Len=0&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Then I analyzed the whole packet capture to find the number of open connections at any time as well as the frequency and duration of these lingering TCP connections. &lt;em&gt;See &lt;a href=&quot;#appendix-analyzing-concurrent-tcp-connections&quot;&gt;Analyzing concurrent TCP connections&lt;/a&gt; for details on how.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;What we find is that the number of open connections at times surge way past 20 and that coincides with increased latency. In addition unused connections stay open for a problematic long time (5-10 seconds).&lt;/p&gt;
&lt;p&gt;Since the closing of connections is initiated by CoreDNS. Is there any way we can make CoreDNS close connections faster or limit the number of connections it uses?&lt;/p&gt;
&lt;p&gt;NodeLocal DNSCache (in GKE at least) is configured (&lt;code&gt;kubectl get configmap --namespace kube-system node-local-dns -o yaml | yq .data.Corefile&lt;/code&gt;) with:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cluster.local:53 {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  forward . __PILLAR__CLUSTER__DNS__ {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      force_tcp&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      expire 1s&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;So it is actually configured to “expire” connections after 1 second.&lt;/p&gt;
&lt;p&gt;I’m not well versed in the CoreDNS codebase but it seems expiring connections is &lt;a href=&quot;https://github.com/coredns/coredns/blob/master/plugin/pkg/proxy/persistent.go#L48&quot;&gt;handled by a ticker&lt;/a&gt;. A ticker in go sends a signal every “tick” that can be used to trigger events for example. This means connections aren’t closed once they pass the 1 second mark, but the cleanup process runs every 1 second and then purges expired connections. So a connection can be idle for almost 2x the time (2 seconds in our case) before being cleaned up.&lt;/p&gt;
&lt;p&gt;I still can’t explain why we see connections lingering on for 5-10 seconds though.&lt;/p&gt;
&lt;p&gt;There are no configuration options indicating that it’s possible to limit the number of connections. Nor anything in the code that would suggest it’s a possibility. Adding it is probably not done in a day either as it would require making decisions on trade-offs about how to handle excess traffic volume. How much to queue, for how long. How to avoid the queue eating too much memory. What to do when the queue is full and so on and so on.&lt;/p&gt;
&lt;p&gt;At this point I feel we have a pretty good grasp of how all of this conspires to cause the problems we observe. But unfortunately I can’t see any permanent solution that does not require modifying CoreDNS and ideally &lt;code&gt;dnsmasq&lt;/code&gt; code itself. At the moment we don’t have the capacity to create and push through such changes upstream. If I could snap my fingers and magically get some new features they would be:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;dnsmasq&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Make &lt;code&gt;MAX_PROCS&lt;/code&gt; configurable.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;CoreDNS&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Configurable max TCP connection pool size. Metrics on pool size and usage.&lt;/li&gt;
&lt;li&gt;Configurable queue size for requests waiting for a new connection from the pool. Metrics on queue capacity (max) and current size for tuning.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But until that becomes a reality I think the best option is to avoid using NodeLocal DNSCache in combination with &lt;code&gt;kube-dns&lt;/code&gt;, and instead replace &lt;code&gt;kube-dns&lt;/code&gt; with Cloud DNS.&lt;/p&gt;
&lt;h2 id=&quot;appendices&quot;&gt;Appendices&lt;/h2&gt;
&lt;p&gt;&lt;a id=&quot;appendix-enabling-additional-envoy-metrics&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;appendix---enabling-additional-envoy-metrics&quot;&gt;Appendix - Enabling additional envoy metrics&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;See &lt;a href=&quot;https://istio.io/latest/docs/concepts/observability/#proxy-level-metrics&quot;&gt;istio.io&lt;/a&gt; for some high level info on proxy-level metrics and &lt;a href=&quot;https://www.envoyproxy.io/docs/envoy/latest/configuration/upstream/cluster_manager/cluster_stats&quot;&gt;envoyproxy.io&lt;/a&gt; for complete list of available metrics.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Most likely we are interested in the &lt;code&gt;upstream_cx_*&lt;/code&gt; and &lt;code&gt;upstream_rq_*&lt;/code&gt; metrics.&lt;/p&gt;
&lt;p&gt;By default metrics are only gathered and exposed for the &lt;code&gt;xds-grpc&lt;/code&gt; cluster, and the full stats name looks like &lt;code&gt;cluster.xds-grpc.upstream_cx_total&lt;/code&gt;. The &lt;code&gt;xds-grpc&lt;/code&gt; cluster I assume is metrics for traffic between the &lt;code&gt;istio-proxy&lt;/code&gt; containers in each Pod and the central istio services (&lt;code&gt;istiod&lt;/code&gt;) used for configuration management.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;A cluster in envoy is a grouping of backends and endpoints. Typically each Kubernetes &lt;code&gt;Service&lt;/code&gt; will be it’s own cluster. There are also separate clusters named &lt;code&gt;BlackHoleCluster&lt;/code&gt;, &lt;code&gt;InboundPassthroughClusterIpv4&lt;/code&gt; and &lt;code&gt;PassthroughCluster&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;When enabling metrics for other services (or clusters as they’re also called) they look like&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;cluster.outbound|80||my-service.my-namespace.svc.cluster.local.upstream_cx_total&lt;/code&gt; for outgoing traffic and&lt;/li&gt;
&lt;li&gt;&lt;code&gt;cluster.inbound|8080||.upstream_cx_total&lt;/code&gt; for incoming traffic.&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;Beware that traffic to for example api.signicat.com/some-service will be mapped to the internal Kubernetes Service DNS name in the metric, like some-service.some-team.svc.cluster.local.&lt;/p&gt;
&lt;p&gt;Ports are also mapped. For outgoing traffic it will be the port in the Service. While for incoming traffic it will be the port the Pod is actually listening on, and not the one mapped in the Service.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id=&quot;enabling-additional-envoy-metrics&quot;&gt;Enabling additional envoy metrics&lt;/h4&gt;
&lt;p&gt;istio.io documents &lt;a href=&quot;https://istio.io/latest/docs/ops/configuration/telemetry/envoy-stats/&quot;&gt;how to enable additional envoy metrics&lt;/a&gt; both globally in the mesh by configuring the &lt;code&gt;IstioOperator&lt;/code&gt; object as well as on a per Deployment/Pod using the &lt;code&gt;proxy.istio.io/config&lt;/code&gt; annotation.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;WARNING: Be very careful when enabling additional metrics as they have a tendency to expose orders of magnitude more time-series than you might expect.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;I apparently managed to get Prometheus OOMKilled even with my fairly limited testing on two deployments.&lt;/p&gt;
&lt;p&gt;If you happen to have &lt;a href=&quot;https://victoriametrics.com/&quot;&gt;VictoriaMetrics&lt;/a&gt; in your monitoring stack you can monitor cardinality (the number of unique time-series, which is the thing that usually breaks a time-series database) in the VM UI:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl -n metrics port-forward services/victoria-metrics-cluster-vmselect 8481:8481&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And going to &lt;a href=&quot;http://localhost:8481/select/0/vmui/?#/cardinality&quot;&gt;http://localhost:8481/select/0/vmui/?#/cardinality&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id=&quot;enable-additional-metrics-on-a-deployment-or-pod&quot;&gt;Enable additional metrics on a Deployment or Pod&lt;/h4&gt;
&lt;p&gt;To enable additional metrics on a Deployment or Pod, add the following annotations:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;yaml&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;metadata&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;  annotations&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;    proxy.istio.io/config&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#D73A49;--shiki-dark:#F97583&quot;&gt;|-&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;      proxyStatsMatcher:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;        inclusionSuffixes:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          - &quot;upstream_rq_timeout&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          - &quot;upstream_rq_retry&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          - &quot;upstream_rq_retry_limit_exceeded&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          - &quot;upstream_rq_retry_success&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          - &quot;upstream_rq_retry_overflow&quot;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Remember to add this to the &lt;strong&gt;Pod&lt;/strong&gt; metadata (&lt;code&gt;spec.template.metadata&lt;/code&gt;) if adding to a Deployment.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;This results in a new time-series being exposed for each Service and Port (Cluster) that envoy has configured.&lt;/p&gt;
&lt;p&gt;In our case we have ~500 services in our dev cluster and enabling these specific metrics adds ~3.800 new time-series for &lt;em&gt;each Pod&lt;/em&gt; in the Deployment we added it to. The test Deployment I’m playing with has 6 Pods so ~23.000 new time-series from adding 5 additional metrics to 1 Deployment!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Another option is to use regex to enable additional metrics:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;yaml&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;          inclusionRegexps&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;          - &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;&quot;.*upstream_rq_.*&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;          - &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;&quot;.*upstream_cx_.*&quot;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;But again, enabling these ~50 metrics on this specific Deployment will result in ~250.000 new time-series.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Be especially wary of metrics envoy exposes as histograms, such as &lt;code&gt;upstream_cx_connect_ms&lt;/code&gt; and &lt;code&gt;upstream_cx_length_ms&lt;/code&gt; as they result in many _bucket time-series. During my testing this resulted in 6-7 million new time-series in total.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It’s possible to specify only the connections we are interested in, which makes it manageable, cardinality wise. For example to only gather metrics from Services A through F in their corresponding namespaces:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;yaml&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;    inclusionRegexps&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;      - &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;&quot;.*(svc-a.ns-a.svc|svc-b.ns-b.svc|svc-c.ns-c.svc|svc-d.ns-d.svc|svc-e.ns-e.svc|svc-f.ns-f.svc).*.upstream_rq.*&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;      - &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;&quot;.*(svc-a.ns-a.svc|svc-b.ns-b.svc|svc-c.ns-c.svc|svc-d.ns-d.svc|svc-e.ns-e.svc|svc-f.ns-f.svc).*.upstream_cx.*&quot;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h4 id=&quot;enable-additional-metrics-globally-please-dont&quot;&gt;Enable additional metrics globally (please don’t!)&lt;/h4&gt;
&lt;p&gt;You can also enable additional metrics globally across the mesh. But that’s probably a very bad idea if you are running at any sort of scale. I estimated that enabling &lt;code&gt;.*upstream_rq_.*&lt;/code&gt; and &lt;code&gt;.*upstream_cx_.*&lt;/code&gt; in our dev cluster would result in 50M additional time-series at a minimum. Or 5-10x our current Prometheus usage.&lt;/p&gt;
&lt;p&gt;If you are using the old &lt;a href=&quot;https://istio.io/latest/docs/setup/install/operator/&quot;&gt;Istio Operator&lt;/a&gt; way of installing and managing istio it should be enough to update the &lt;code&gt;IstioOperator&lt;/code&gt; object in the cluster. If you are using &lt;code&gt;istioctl&lt;/code&gt; (&lt;a href=&quot;https://istio.io/latest/about/faq/#install-method-selection&quot;&gt;recommended&lt;/a&gt;) you must update the source &lt;code&gt;IstioOperator&lt;/code&gt; manifest that is being fed to &lt;code&gt;istioctl&lt;/code&gt; and run the appropriate &lt;code&gt;istioctl&lt;/code&gt; commands again to update. Note that this also creates an &lt;code&gt;IstioOperator&lt;/code&gt; object in the cluster with whatever config is used. But in this case it’s never used for anything other than reference. So updating the &lt;code&gt;IstioOperator&lt;/code&gt; object in a cluster if managing istio with &lt;code&gt;istioctl&lt;/code&gt; does nothing.&lt;/p&gt;
&lt;h4 id=&quot;viewing-stats-and-metrics&quot;&gt;Viewing stats and metrics&lt;/h4&gt;
&lt;p&gt;Metrics should start becoming available in Prometheus with names like &lt;code&gt;envoy_cluster_upstream_cx_total&lt;/code&gt;. Note that by default you’ll already see metrics from the &lt;code&gt;xds-grpc&lt;/code&gt; cluster.&lt;/p&gt;
&lt;p&gt;You can also get the &lt;a href=&quot;https://istio.io/latest/docs/ops/configuration/telemetry/envoy-stats/&quot;&gt;stats directly from a sidecar&lt;/a&gt;. Either by querying the &lt;code&gt;pilot-agent&lt;/code&gt; directly:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl exec -n my-namespace my-pod -c istio-proxy -- pilot-agent request GET stats&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Or querying the metrics endpoint exposed to Prometheus:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl exec -n my-namespace my-pod -c istio-proxy -- curl -sS &apos;localhost:15000/stats/prometheus&apos;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Note that these will not give you any additional metrics compared to those exposed to Prometheus. A metric that is not enabled will not show up, and all metrics that are enabled are also automatically exposed to Prometheus.&lt;/p&gt;
&lt;h4 id=&quot;notes-on-timeout-and-retry-metrics&quot;&gt;Notes on timeout and retry metrics&lt;/h4&gt;
&lt;p&gt;If istio is configured to be involved with timeouts and retries that is configured on the &lt;a href=&quot;https://istio.io/latest/docs/reference/config/networking/virtual-service/#Destination&quot;&gt;VirtualService&lt;/a&gt; level.&lt;/p&gt;
&lt;p&gt;That means it will only take effect if using &lt;code&gt;api.signicat.com/some-service&lt;/code&gt; (a route in a Istio &lt;code&gt;VirtualService&lt;/code&gt;) and not &lt;code&gt;some-service.some-namespace.svc.cluster.local&lt;/code&gt; (Kubernetes &lt;code&gt;Service&lt;/code&gt;).&lt;/p&gt;
&lt;p&gt;Istio will only count timeouts and retries for requests to a &lt;code&gt;VirtualService&lt;/code&gt; route that has timeouts (and optionally retries) configured.&lt;/p&gt;
&lt;p&gt;Additionally, the timeouts and retries seems to be enforced, and counted, in the Source istio-proxy (envoy), and not the Target.&lt;/p&gt;
&lt;h4 id=&quot;further-work&quot;&gt;Further work&lt;/h4&gt;
&lt;p&gt;It would be beneficial to be able to have envoy only show/export metrics that are non-zero. Since there is usually only a very small set of all possible Service-to-Service pairs that will actually regularly have traffic.&lt;/p&gt;
&lt;p&gt;It’s possible to customize metrics using the Telemetry API (&lt;a href=&quot;https://istio.io/latest/docs/tasks/observability/metrics/telemetry-api/&quot;&gt;Customizing Istio Metrics with Telemetry API&lt;/a&gt;) but it seems limited to only working with metrics and their dimensions. Not the time-series values.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://istio.io/latest/docs/reference/config/proxy_extensions/wasm-plugin/&quot;&gt;WASM plugins&lt;/a&gt; are probably not a good fit either and are experimental and causes a severe CPU and memory penalty.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Edit to add in 2024&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;We know today that the reason for so many time-series (reporting zero), as well as elevated memory usage in &lt;code&gt;istio-proxy&lt;/code&gt;, is because we did not filter exposed Services between namespaces causing every &lt;code&gt;istio-proxy&lt;/code&gt; to keep track of every pair of possible workloads in the whole cluster.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;appendix-overview-of-dns-on-gke&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;appendix---overview-of-dns-on-gke&quot;&gt;Appendix - Overview of DNS on GKE&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;kube-dns and CoreDNS confusion&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The main difference between DNS on upstream Kubernetes and GKE is that Kubernetes &lt;a href=&quot;https://kubernetes.io/blog/2018/07/10/coredns-ga-for-kubernetes-cluster-dns/&quot;&gt;switched to&lt;/a&gt; CoreDNS 5 years ago while GKE &lt;a href=&quot;https://cloud.google.com/kubernetes-engine/docs/how-to/kube-dns&quot;&gt;still uses&lt;/a&gt; the old &lt;code&gt;kube-dns&lt;/code&gt;. It’s &lt;a href=&quot;https://cloud.google.com/knowledge/kb/how-to-run-coredns-on-kubernetes-engine-000004698&quot;&gt;possible to add CoreDNS but it’s not possible to remove or disable&lt;/a&gt; &lt;code&gt;kube-dns&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;In upstream K8s, &lt;a href=&quot;https://kubernetes.io/blog/2018/07/10/coredns-ga-for-kubernetes-cluster-dns/&quot;&gt;CoreDNS reached GA back in 2018 in K8s 1.11&lt;/a&gt; and &lt;code&gt;kube-dns&lt;/code&gt; was &lt;a href=&quot;https://kubernetes.io/docs/tasks/administer-cluster/coredns/#migrating-to-coredns&quot;&gt;removed from kubeadm in K8s 1.21&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;However for &lt;a href=&quot;https://github.com/coredns/deployment/issues/116&quot;&gt;backwards compatibility&lt;/a&gt; CoreDNS &lt;a href=&quot;https://github.com/coredns/deployment/blob/master/kubernetes/coredns.yaml.sed&quot;&gt;still uses&lt;/a&gt; the name &lt;code&gt;kube-dns&lt;/code&gt;, which makes things confusing for sure!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Diagram showing DNS on GKE&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./DNS-on-GKE.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;Part of the magic of Kubernetes is it’s DNS-based Service Discovery. Lets say we deploy an application that has a &lt;code&gt;Service&lt;/code&gt; named &lt;code&gt;helloworld&lt;/code&gt; to the namespace &lt;code&gt;team-a&lt;/code&gt;. Other applications in the same namespace are able to connect to that service by calling for example &lt;code&gt;http://helloworld&lt;/code&gt;. This makes it easy to build and configure a group of microservices that talk to each other without needing to know about namespaces or FQDNs. Applications in another namespace can also call that application using &lt;code&gt;http://helloworld.team-a&lt;/code&gt;. (&lt;a href=&quot;#footnote-d&quot;&gt;Footnote D&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;This magic is achieved using the standard DNS &lt;code&gt;search&lt;/code&gt; domain list feature. On Linux the &lt;code&gt;/etc/resolv.conf&lt;/code&gt; &lt;a href=&quot;https://man7.org/linux/man-pages/man5/resolv.conf.5.html&quot;&gt;file&lt;/a&gt; defines which DNS servers to use. It also includes a &lt;code&gt;search&lt;/code&gt; option that works together with the &lt;code&gt;ndots&lt;/code&gt; option. You can see this by executing &lt;code&gt;cat /etc/resolv.conf&lt;/code&gt; in any &lt;code&gt;Pod&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;If a domain name contains fewer “dots” (“.”) than &lt;code&gt;ndots&lt;/code&gt; is set to, the DNS resolver will first try the hostname with each of the appended domains in &lt;code&gt;search&lt;/code&gt;. In Kubernetes (and GKE) ndots is by default set to 5.&lt;/p&gt;
&lt;p&gt;In vanilla Kubernetes the search domains are&lt;code&gt; &amp;lt;namespace&amp;gt;.svc.cluster.local svc.cluster.local cluster.local&lt;/code&gt; while on GKE it’s &lt;code&gt;&amp;lt;namespace&amp;gt;.svc.cluster.local svc.cluster.local cluster.local &amp;lt;gcp-zone&amp;gt;.c.technology-dev-platform.internal c.&amp;lt;gcp-project&amp;gt;.internal google.internal&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;This means if we call &lt;code&gt;https://api.signicat.com&lt;/code&gt; (2 dots) it will first try to resolve &lt;code&gt;api.signicat.com.some-namespace.svc.cluster.local&lt;/code&gt; and so on through the whole list, sequentally, before finally trying &lt;code&gt;api.signicat.com&lt;/code&gt; and actually succeeding.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;See &lt;a href=&quot;#appendix-reducing-dns-lookups&quot;&gt;Reducing DNS lookups in Kubernetes and GKE&lt;/a&gt; for thoughts on avoiding some of this.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Then the requests are sent to the &lt;code&gt;dnsmasq&lt;/code&gt; container of a &lt;code&gt;kube-dns&lt;/code&gt; Pod. &lt;code&gt;dnsmasq&lt;/code&gt; is configured by a combination of command-line arguments defined in the &lt;code&gt;kube-dns&lt;/code&gt; Deployment and data from the &lt;code&gt;kube-dns&lt;/code&gt; &lt;a href=&quot;https://cloud.google.com/kubernetes-engine/docs/how-to/kube-dns&quot;&gt;ConfigMap&lt;/a&gt;. We have not changed the default &lt;code&gt;kube-dns&lt;/code&gt; ConfigMap and this results in &lt;code&gt;dnsmasq&lt;/code&gt; running with these arguments:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;/usr/sbin/dnsmasq -k \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --cache-size=1000 \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --no-negcache \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --dns-forward-max=1500 \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --log-facility=- \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --server=/cluster.local/127.0.0.1#10053 \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --server=/in-addr.arpa/127.0.0.1#10053 \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --server=/ip6.arpa/127.0.0.1#10053 \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --max-ttl=30 \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --max-cache-ttl=30 \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --server /internal/169.254.169.254 \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --server 8.8.8.8 \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --server 8.8.4.4 \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --no-resolv&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This means all lookups for &lt;code&gt;cluster.local&lt;/code&gt;, &lt;code&gt;in-addr.arpa&lt;/code&gt; and &lt;code&gt;ip6.arpa&lt;/code&gt; will be sent to &lt;code&gt;127.0.0.1:10053&lt;/code&gt;, which is the &lt;code&gt;kube-dns&lt;/code&gt; container. Lookups for &lt;code&gt;internal&lt;/code&gt; are sent to &lt;code&gt;169.254.169.254&lt;/code&gt;. All others are sent to &lt;code&gt;8.8.8.8&lt;/code&gt; and &lt;code&gt;8.8.4.4&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;In addition &lt;code&gt;--no-negcache&lt;/code&gt; disables caching for lookups that were not found (&lt;code&gt;NXDOMAIN&lt;/code&gt;). This is particularly interesting since that means when looking up for example &lt;code&gt;api.signicat.com&lt;/code&gt; and the domains in the &lt;code&gt;search&lt;/code&gt; list are first tried, they will all result in &lt;code&gt;NXDOMAIN&lt;/code&gt; but &lt;code&gt;dnsmasq&lt;/code&gt; will not cache those results but send them on to &lt;code&gt;kube-dns&lt;/code&gt;. &lt;em&gt;Every&lt;/em&gt;. &lt;em&gt;Time&lt;/em&gt;. This may severely reduce the usefulness of the cache and create a constant volume of traffic to &lt;code&gt;kube-dns&lt;/code&gt; itself.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;It’s also worth noting that Google Public DNS servers have a &lt;a href=&quot;https://developers.google.com/speed/public-dns/docs/isp&quot;&gt;default rate limit of 1500 QPS per IP&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;appendix-reducing-dns-lookups&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;appendix---reducing-dns-lookups-in-kubernetes-and-gke&quot;&gt;Appendix - Reducing DNS lookups in Kubernetes and GKE&lt;/h3&gt;
&lt;p&gt;There are two reasons for the 7 DNS lookups instead of the ideal 1. First is the &lt;code&gt;ndots&lt;/code&gt; which tells the DNS resolver in the container if a domain name has fewer dots than this, it will first try looking up the name by appending each entry in the &lt;code&gt;search&lt;/code&gt; list sequentially. For Kubernetes &lt;code&gt;ndots&lt;/code&gt; is set to 5 (why is &lt;a href=&quot;https://github.com/kubernetes/kubernetes/issues/33554#issuecomment-266251056&quot;&gt;explained by Tim Hockin here&lt;/a&gt;). So &lt;code&gt;api.signicat.com&lt;/code&gt; only has 2 dots, and hence will first go through the list of search domains.&lt;/p&gt;
&lt;p&gt;Secondly GKE adds 3 extra search domains in addition to the 3 standard ones in Kubernetes. Bringing the total to 6 before doing the “proper” DNS lookup.&lt;/p&gt;
&lt;p&gt;The specific assumption we are deviating from leading to problems in our case is “We could mitigate some of the perf penalties by always trying names as upstream FQDNs first, but that means that all intra-cluster lookups get slower. Which do we expect more frequently? I’ll argue intra-cluster names, if only because the TTL is so low.”&lt;/p&gt;
&lt;p&gt;An option that we’re currently exploring is using a trailing dot in the FQDN (&lt;code&gt;api.signicat.com.&lt;/code&gt;) when calling other services. This explicitly tells DNS that this is a FQDN and should not search through the search domain lists for a hit first. This seems to work on some of our services but not all. Indicating that there isn’t any inherent problems doing this with regards to istio or other infrastructure. But there may be additional changes needed on some services to support this. I’m suspecting certain web application frameworks in certain languages not handling this well out of the box.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;appendix-overview-of-gke-dns-with-nodelocal-dnscache&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;appendix---overview-of-gke-dns-with-nodelocal-dnscache&quot;&gt;Appendix - Overview of GKE DNS with NodeLocal DNSCache&lt;/h3&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Diagram showing DNS on GKE with NodeLocal DNSCache enabled&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./DNS-on-GKE-with-Node-Local-DNS-Cache.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://kubernetes.io/docs/tasks/administer-cluster/nodelocaldns/&quot;&gt;NodeLocal DNSCache&lt;/a&gt; is a feature of Kubernetes that primarily aims to improve performance, scale and latency by:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Potentially not traversing the network to another node&lt;/li&gt;
&lt;li&gt;Skip iptables DNAT which sometimes caused problems&lt;/li&gt;
&lt;li&gt;Upgrading connections from UDP to TCP which should reduce latency in case of dropped packets&lt;/li&gt;
&lt;li&gt;Enabling negative caching&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;NodeLocal DNSCache is &lt;a href=&quot;https://cloud.google.com/kubernetes-engine/docs/how-to/nodelocal-dns-cache&quot;&gt;available as a GKE add-on&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;NodeLocal DNSCache is deployed as a &lt;code&gt;DaemonSet&lt;/code&gt; named &lt;code&gt;node-local-dns&lt;/code&gt; in &lt;code&gt;kube-system&lt;/code&gt; namespace. One &lt;code&gt;node-local-dns-x&lt;/code&gt; Pod is created on each node in the cluster.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;node-local-dns-x&lt;/code&gt; Pods run CoreDNS and are configured by the &lt;code&gt;node-local-dns&lt;/code&gt; ConfigMap (&lt;code&gt;kubectl get cm node-local-dns -o yaml|yq .data.Corefile&lt;/code&gt;).&lt;/p&gt;
&lt;p&gt;The configuration is similar to &lt;code&gt;dnsmasq&lt;/code&gt; in that requests for &lt;code&gt;cluster.local&lt;/code&gt;, &lt;code&gt;in-addr.arpa&lt;/code&gt; and &lt;code&gt;ip6.arpa&lt;/code&gt; are sent to &lt;code&gt;kube-dns&lt;/code&gt; while the rest are sent to &lt;code&gt;8.8.8.8&lt;/code&gt; and &lt;code&gt;8.8.4.4&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;It binds to (re-uses) the same IP as the &lt;code&gt;kube-dns&lt;/code&gt; Service. That way all requests from Pods on the node towards &lt;code&gt;kube-dns&lt;/code&gt; will actually be handled by &lt;code&gt;node-local-dns&lt;/code&gt; instead. And to actually communicate with &lt;code&gt;kube-dns&lt;/code&gt; another Service named &lt;code&gt;kube-dns-upstream&lt;/code&gt; is created that is a clone of the &lt;code&gt;kube-dns&lt;/code&gt; Service but with a different IP.&lt;/p&gt;
&lt;p&gt;Even though &lt;code&gt;node-local-dns&lt;/code&gt; uses CoreDNS and &lt;a href=&quot;https://github.com/prometheus-community/helm-charts/tree/main/charts/kube-prometheus-stack&quot;&gt;kube-prometheus-stack&lt;/a&gt; has support for scraping CoreDNS it won’t necessarily work for &lt;code&gt;node-local-dns&lt;/code&gt;. See the next appendix for how to scrape metrics from &lt;code&gt;node-local-dns&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;appendix-enabling-node-local-dns-metrics&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;appendix---enabling-node-local-dns-metrics&quot;&gt;Appendix - Enabling node-local-dns metrics&lt;/h3&gt;
&lt;p&gt;We wanted to look at some metrics from node-local-dns, but found that we didn’t have any!&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;node-local-dns&lt;/code&gt; pods do have annotations for scraping metrics:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;yaml&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;metadata&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;  annotations&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;    prometheus.io/port&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;&quot;9253&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#22863A;--shiki-dark:#85E89D&quot;&gt;    prometheus.io/scrape&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;: &lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;&quot;true&quot;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;But in our Prometheus we don’t use the &lt;code&gt;kubernetes-pods&lt;/code&gt; &lt;a href=&quot;https://github.com/prometheus-community/helm-charts/blob/main/charts/prometheus/values.yaml#L998&quot;&gt;scrape job config from the prometheus example&lt;/a&gt;. meaning that these targets will not be discovered or scraped by prometheus. (And we don’t want to add and allow for scraping this way).&lt;/p&gt;
&lt;p&gt;The &lt;a href=&quot;https://github.com/prometheus-community/helm-charts/tree/main/charts/kube-prometheus-stack&quot;&gt;kube-prometheus-stack helm chart&lt;/a&gt; comes with both a &lt;a href=&quot;https://github.com/prometheus-community/helm-charts/blob/main/charts/kube-prometheus-stack/templates/exporters/core-dns/service.yaml&quot;&gt;Service&lt;/a&gt; and &lt;a href=&quot;https://github.com/prometheus-community/helm-charts/blob/main/charts/kube-prometheus-stack/templates/exporters/core-dns/servicemonitor.yaml&quot;&gt;ServiceMonitor&lt;/a&gt; to scrape metrics from CoreDNS (which NodeLocal DNSCache uses).&lt;/p&gt;
&lt;p&gt;However the Service and ServiceMonitor uses a hardcoded selector of &lt;code&gt;k8s-app: kube-dns&lt;/code&gt; while &lt;code&gt;node-local-dns&lt;/code&gt; Pods have &lt;code&gt;k8s-app: node-local-dns&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;We made some changes to these and deployed them to the cluster and now started getting metrics in the &lt;code&gt;coredns_dns_*&lt;/code&gt; timeserieses.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;You can add our updated Service and ServiceMonitor like this:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl apply -f -n kube-system https://raw.githubusercontent.com/signicat/blog-attachements/main/2023-gke-node-local-dns-cache/files/node-local-dns-metrics-service.yaml&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl apply -f -n metrics https://raw.githubusercontent.com/signicat/blog-attachements/main/2023-gke-node-local-dns-cache/files/node-local-dns-servicemonitor.yaml&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And metrics should start flowing within a couple of minutes.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In new versions of the &lt;a href=&quot;https://github.com/prometheus-community/helm-charts/blob/main/charts/kube-prometheus-stack/templates/grafana/dashboards-1.14/k8s-coredns.yaml&quot;&gt;CoreDNS dashboard&lt;/a&gt; that comes with &lt;code&gt;kube-prometheus-stack&lt;/code&gt; you should be able to select the &lt;code&gt;node-local-dns&lt;/code&gt; job.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I added the &lt;code&gt;job&lt;/code&gt; template variable to the dashboard in &lt;a href=&quot;https://github.com/prometheus-community/helm-charts/pull/3798&quot;&gt;this PR&lt;/a&gt;, which may not have made it into a new version of the &lt;code&gt;kube-prometheus-stack&lt;/code&gt; chart yet. In the mean time you can use &lt;a href=&quot;https://github.com/signicat/blog-attachements/blob/main/2023-gke-node-local-dns-cache/files/grafana-dashboard-coredns.json&quot;&gt;our updated CoreDNS dashboard&lt;/a&gt; which adds the necessary template variable as well as a couple of other improvements.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;appendix-using-dnsperf&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;appendix---using-dnsperf-to-test-dns-performance&quot;&gt;Appendix - Using dnsperf to test DNS performance&lt;/h3&gt;
&lt;p&gt;Trying to tease out more information on the problem we run some DNS load testing using &lt;a href=&quot;https://linux.die.net/man/1/dnsperf&quot;&gt;dnsperf&lt;/a&gt;.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;An alternative to &lt;code&gt;dnsperf&lt;/code&gt; is &lt;a href=&quot;https://linux.die.net/man/1/resperf&quot;&gt;resperf&lt;/a&gt; which is “a companion tool to dnsperf” designed for testing resolution performance of a caching DNS server. Whereas &lt;code&gt;dnsperf&lt;/code&gt; is primarily meant to test authoritative DNS servers. Since &lt;code&gt;resperf&lt;/code&gt; is more complicated to work with, and we want to test at a fairly low volume, we assume that &lt;code&gt;dnsperf&lt;/code&gt; is good enough for now.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;We spin up a new pod with Ubuntu to run dnsperf from&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl run dnsperf -n default --image=ubuntu:22.04 -i --tty --restart=Never&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And install dnsperf&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apt-get update &amp;amp;&amp;amp; apt-get install -y dnsperf&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;We’ll be testing three different destinations.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;kube-dns&lt;/code&gt; Service: First we use the IP of the &lt;code&gt;kube-dns&lt;/code&gt; Service (&lt;code&gt;172.18.0.10&lt;/code&gt;) which will be intercepted by the &lt;code&gt;node-local-dns&lt;/code&gt; Pod on the same node. It will do it’s caching and the traffic is visible in our modified CoreDNS Grafana dashboard for that node.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;kube-dns-upstream&lt;/code&gt; Service: Then we use the IP of the &lt;code&gt;kube-dns-upstream&lt;/code&gt; Service (&lt;code&gt;172.18.165.139&lt;/code&gt;) which is a copy of the &lt;code&gt;kube-dns&lt;/code&gt; Service that &lt;code&gt;node-local-dns&lt;/code&gt; uses when looking up &lt;code&gt;.cluster.local&lt;/code&gt; domains. It has a different IP than &lt;code&gt;kube-dns&lt;/code&gt; so that it won’t be intercepted by &lt;code&gt;node-local-dns&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;One &lt;code&gt;kube-dns&lt;/code&gt; Pod: Then we use the IP of one specific &lt;code&gt;kube-dns&lt;/code&gt; Pod (&lt;code&gt;172.20.14.115&lt;/code&gt;) so that we can observe the behaviour of one instance without any load balancing sprinkling the traffic all over the place.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;code&gt;dnsperf&lt;/code&gt; takes a list of domains to look up in a file that looks like&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;archive.ubuntu.com A&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;test.salesforce.com A&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;storage.googleapis.com A&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;To create a file with the domains we actually look up I do a packet capture on one of the &lt;code&gt;node-local-dns&lt;/code&gt; Pods for half an hour. Open it in Wireshark. Filter by &lt;code&gt;dns.flags.response == 1&lt;/code&gt;. Add DNS -&amp;gt; Queries -&amp;gt; query -&amp;gt; Name as a Column. Export Packet Dissections as CSV. And use Excel and Notepad to get a file of 50.000 DNS names in that format. Upload that file to the newly created &lt;code&gt;dnsperf&lt;/code&gt; Pod. Create a separate file with only &lt;code&gt;cluster.local&lt;/code&gt; domains as well (&lt;code&gt;cat domains-all.txt | grep &quot;cluster.local&quot; &amp;gt; domains-cluster-local.txt&lt;/code&gt;).&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;dnsperf&lt;/code&gt; does not expand/multiply domains by adding the domains from &lt;code&gt;search-path&lt;/code&gt; in &lt;code&gt;/etc/resolve.conf&lt;/code&gt;. The list we extracted from the packet capture already have these additional lookups though so it is representative nonetheless.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Run dnsperf:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;dnsperf -d domains-all.txt -l 60 -Q 100 -S 5 -t 2 -s 172.18.0.10&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Explanation of arguments:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;-d domains-all.txt&lt;/code&gt; - Domain list file.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;-l 60&lt;/code&gt; - Length (duration) of test in seconds.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;-Q 100&lt;/code&gt; - Queries per second.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;-S 5&lt;/code&gt; - Print statistics every 5 seconds.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;-t 2&lt;/code&gt; - Timeout 2 seconds.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;-s 172.18.0.18&lt;/code&gt; - DNS server to query&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Each test has different values for the arguments.&lt;/p&gt;
&lt;p&gt;Results from running a series of tests ranging from 100 Queries per second (QPS) to 1500 QPS:&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Load test results table&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./2023-12-14-table-load-tests.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;These results are from a second round of testing a couple of weeks after the first. On the first round I observed a lot (6%) timing out on just 200 QPS directly towards a &lt;code&gt;kube-dns&lt;/code&gt; Pod. I’m not sure why the results are suddenly much better. Looking at the graphs from &lt;code&gt;node-local-dns&lt;/code&gt; there is a significant improvement overall some days before the second round of testing. We have not done any changes that could explain the sudden improvement. I guess it’s just one of those things…&lt;/p&gt;
&lt;p&gt;CoreDNS did &lt;a href=&quot;https://coredns.io/2018/11/27/cluster-dns-coredns-vs-kube-dns/&quot;&gt;a benchmark&lt;/a&gt; of &lt;code&gt;kube-dns&lt;/code&gt; and &lt;code&gt;CoreDNS&lt;/code&gt; and managed to get ~36.000 QPS on internal names and ~2.200 QPS on external names on &lt;code&gt;kube-dns&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;appendix-network-packet-capture&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;appendix---network-packet-capture-without-dependencies&quot;&gt;Appendix - Network packet capture without dependencies&lt;/h3&gt;
&lt;p&gt;I haven’t done packet capture in a Kubernetes cluster before. First I tried &lt;a href=&quot;https://github.com/eldadru/ksniff&quot;&gt;ksniff&lt;/a&gt; that &lt;a href=&quot;https://anythingsimple.medium.com/how-to-do-network-sniff-for-kubernetes-pod-running-on-gke-fb23d0b63e95&quot;&gt;this blog post describes&lt;/a&gt;. But no dice.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Another method that is much more reliable is running tcpdump in a &lt;a href=&quot;https://kubernetes.io/docs/tasks/debug/debug-application/debug-running-pod/#ephemeral-container&quot;&gt;debug container&lt;/a&gt; as mentioned &lt;a href=&quot;https://downey.io/blog/kubernetes-ephemeral-debug-container-tcpdump/&quot;&gt;here&lt;/a&gt;. He pipes tcpdump from a container directly to wireshark on his own machine. But since I’m using &lt;code&gt;kubectl&lt;/code&gt; etc inside Ubuntu on WSL2 on Windows that is probably going to require much more setup.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Instead I’m opting for just saving packet captures in the container as files and copy them to a directory on my machine accessible by Windows.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Spin up a new container using the ubuntu:22.04 image in the existing node-local-dns-7vqpp Pod. Try to attach to the process-namespace of the node-cache container if possible.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl debug -c debug -n kube-system -it node-local-dns-7vqpp --image=ubuntu:22.04 --target=node-cache&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;As soon as the image is downloaded and container started your console should give you a new shell inside the container where we can start preparing:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Update apt and install tshark. Enter &quot;no&quot; for &quot;Should non-superusers be able to capture packets?&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apt-get update &amp;amp;&amp;amp; apt-get install -y tshark&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Capture 100 packets from eth0 to verify that things are working&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;tshark -i eth0 -n -c 100&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;At first I captured 1 million packets without any filters. That resulted in a 6GB file which may be a bit on the large side when I’m just interested in some DNS lookups. So lets add a capture filter &lt;code&gt;port 53&lt;/code&gt; going forward.&lt;/p&gt;
&lt;p&gt;Capturing DNS packets to and from a &lt;code&gt;node-local-dns&lt;/code&gt; Pod:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# In the shell running in the debug container with tshark installed:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;date; tshark -i eth0 -n -c 1000000 -f &quot;port 53&quot; -w node-local-dns-7vqpp-eth0.pcap&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;After either capturing 100.000 packets or I’m happy I stop the capture with &lt;code&gt;Ctrl+C&lt;/code&gt; and I can download the pcap file in WSL:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# On your local machine:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl cp -c debug kube-system/node-local-dns-7vqpp:node-local-dns-7vqpp-eth0.pcap /mnt/c/Data/node-local-dns-7vqpp-eth0.pcap&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Then browse to &lt;code&gt;C:\Data\&lt;/code&gt; where you can open &lt;code&gt;node-local-dns-7vqpp-eth0.pcap&lt;/code&gt; in &lt;a href=&quot;https://www.wireshark.org/&quot;&gt;Wireshark&lt;/a&gt;.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I print the current time on the Pod just before starting packet capture. Most of the time the correct UTC time will be recorded on the packets. But in case it isn’t and the time starts from 0, I can use that time to adjust the offset at least roughly. Making it easier to correlate packet captures from different Pods.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;To capture on the receiving &lt;code&gt;kube-dns&lt;/code&gt; Pod, both incoming DNS traffic to the &lt;code&gt;dnsmasq&lt;/code&gt; container on port 53 on &lt;code&gt;eth0&lt;/code&gt;, and between &lt;code&gt;dnsmasq&lt;/code&gt; and &lt;code&gt;kube-dns&lt;/code&gt; containers on &lt;code&gt;lo&lt;/code&gt; port 10053:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl debug -c debug -n kube-system -it kube-dns-d7bc86d4c-d2x8p --image=ubuntu:22.04 --target=dnsmasq&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Install tshark as shown above&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Start two packet captures running in the background:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;date; tshark -i eth0 -n -c 1000000 -f &quot;port 53&quot; -w kube-dns-d7bc86d4c-d2x8p-eth0.pcap &amp;amp;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;date; tshark -i lo -n -c 1000000 -f &quot;port 10053&quot; -w kube-dns-d7bc86d4c-d2x8p-lo.pcap &amp;amp;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# To stop these type `fg` in the shell to bring a background process to the foreground and stop it with `Ctrl+C`.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;# Then `fg` and `Ctrl+C` again to stop the other.&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Download the files to my local machine as before:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl cp -c debug kube-system/kube-dns-d7bc86d4c-d2x8p:kube-dns-d7bc86d4c-d2x8p-eth0.pcap /mnt/c/Data/kube-dns-d7bc86d4c-d2x8p-eth0.pcap&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl cp -c debug kube-system/kube-dns-d7bc86d4c-d2x8p:kube-dns-d7bc86d4c-d2x8p-lo.pcap /mnt/c/Data/kube-dns-d7bc86d4c-d2x8p-lo.pcap&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;em&gt;Check out &lt;a href=&quot;#analyzing-dns-problems-based-on-packet-capture&quot;&gt;Analyzing DNS problems based on packet captures&lt;/a&gt; for some tips and tricks on analyzing the packet captures.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;appendix-verbose-logging-on-kube-dns&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;appendix---verbose-logging-on-kube-dns&quot;&gt;Appendix - Verbose logging on &lt;code&gt;kube-dns&lt;/code&gt;&lt;/h3&gt;
&lt;p&gt;The &lt;a href=&quot;https://dnsmasq.org/docs/dnsmasq-man.html&quot;&gt;dnsmasq man page&lt;/a&gt; lists several interesting options such as &lt;code&gt;--log-queries=extra&lt;/code&gt; and &lt;code&gt;--log-debug&lt;/code&gt;!&lt;/p&gt;
&lt;p&gt;But it’s not possible to make changes to the &lt;code&gt;kube-dns&lt;/code&gt; Deployment since it’s managed by GKE. Any changes you make will be reverted immediately.&lt;/p&gt;
&lt;p&gt;Instead we take the existing manifests for &lt;code&gt;kube-dns&lt;/code&gt;, modify them and create a parallel deployment that we control:&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/signicat/blog-attachements/blob/main/2023-gke-node-local-dns-cache/files/kube-dns-deployment.yaml&quot;&gt;kube-dns-deployment.yaml&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Deployment named &lt;code&gt;kube-dns-debug&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;A couple of annotations commented out that would otherwise make the Deployment be instantly removed.&lt;/li&gt;
&lt;li&gt;Keep existing &lt;code&gt;k8s-app: kube-dns&lt;/code&gt; label on the Pods so they will receive traffic for the &lt;code&gt;kube-dns-upstream&lt;/code&gt; Service.&lt;/li&gt;
&lt;li&gt;An additional &lt;code&gt;reason: debug&lt;/code&gt; label on the Pods so we can target them specifically.&lt;/li&gt;
&lt;li&gt;Additional logging enabled with &lt;code&gt;--log-queries=extra&lt;/code&gt;, &lt;code&gt;--log-debug&lt;/code&gt; and &lt;code&gt;--log-async=25&lt;/code&gt;. I enabled these one at a time but the log volume with everything enabled isn’t overwhelming.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/signicat/blog-attachements/blob/main/2023-gke-node-local-dns-cache/files/kube-dns-debug-service.yaml&quot;&gt;kube-dns-debug-service.yaml&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Service named &lt;code&gt;kube-dns-debug&lt;/code&gt; with port udp+tcp/53 targeting only the Pods with the added &lt;code&gt;reason: debug&lt;/code&gt; label. I used this service to run load tests only towards these Pods.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/signicat/blog-attachements/blob/main/2023-gke-node-local-dns-cache/files/kube-dns-metrics-service.yaml&quot;&gt;kube-dns-metrics-service.yaml&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Service named &lt;code&gt;kube-dns-debug-metrics&lt;/code&gt; with port tcp/10054 and tcp/10055 targeting the same Pods as above but for exposing metrics.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/signicat/blog-attachements/blob/main/2023-gke-node-local-dns-cache/files/kube-dns-servicemonitor.yaml&quot;&gt;kube-dns-servicemonitor.yaml&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;ServiceMonitor named &lt;code&gt;kube-dns-debug&lt;/code&gt; that targets the &lt;code&gt;kube-dns-debug-metrics&lt;/code&gt; Service above.&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;Strictly speaking the &lt;code&gt;kube-dns-debug&lt;/code&gt; Service, &lt;code&gt;kube-dns-debug-metrics&lt;/code&gt; Service and &lt;code&gt;kube-dns-debug&lt;/code&gt; ServiceMonitor isn’t necessary if using the &lt;code&gt;k8s-app: kube-dns&lt;/code&gt; label on the new Pods. But it makes it possible to separate the two deployments completely by using another label on the Pods such as &lt;code&gt;k8s-app: kube-dns-debug&lt;/code&gt; and thus avoid for example load tests affecting real cluster DNS traffic.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Now some DNS requests should arrive at the &lt;code&gt;kube-dns-debug&lt;/code&gt; Pods with additional logging enabled.&lt;/p&gt;
&lt;p&gt;It’s also possible to force all DNS traffic to this Pod by manually adding the &lt;code&gt;reason: debug&lt;/code&gt; label as a selector on the &lt;code&gt;kube-dns-upstream&lt;/code&gt; Service. It will stay that way for “a while” (hours, maybe days) before being reverted. Plenty of time to play around at least.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;appendix-analyzing-dnsmasq-logs&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;appendix---analyzing-dnsmasq-logs&quot;&gt;Appendix - Analyzing &lt;code&gt;dnsmasq&lt;/code&gt; logs&lt;/h3&gt;
&lt;p&gt;Once we have increased the log verbosity of &lt;code&gt;dnsmasq&lt;/code&gt; we can see if there’s anything to learn there.&lt;/p&gt;
&lt;p&gt;First I used &lt;a href=&quot;https://github.com/signicat/blog-attachements/blob/main/2023-gke-node-local-dns-cache/files/grafana-dashboard-coredns.json&quot;&gt;our updated CoreDNS dashboard&lt;/a&gt; to identify a time interval where we observed latency spikes. Then using loki, our centralized log storage, I download about 5000 lines of logs spanning about 50 seconds worth of logs. (Our loki is limited to loading 5000 lines, therefore it’s important to try to narrow down and find logs where we are actually experiencing issues).&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;2023-10-10T10:02:36+02:00 I1010 08:02:36.571446   1 nanny.go:146] dnsmasq[4811]: 315207 10.0.0.83/56382 forwarded metadata.google.internal.cluster.local to 127.0.0.1#10053&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;2023-10-10T10:02:36+02:00 I1010 08:02:36.571457   1 nanny.go:146] dnsmasq[4811]: 315207 10.0.0.83/56382 reply metadata.google.internal.cluster.local is NXDOMAIN&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;2023-10-10T10:02:36+02:00 I1010 08:02:36.571488   1 nanny.go:146] dnsmasq[4810]: 315107 10.0.0.83/24595 forwarded metadata.google.internal.cluster.local to 127.0.0.1#10053&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;2023-10-10T10:02:36+02:00 I1010 08:02:36.571495   1 nanny.go:146] dnsmasq[4810]: 315107 10.0.0.83/24595 reply metadata.google.internal.cluster.local is NXDOMAIN&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;2023-10-10T10:02:36+02:00 I1010 08:02:36.575810   1 nanny.go:146] dnsmasq[4810]: 315108 10.0.0.83/24595 query[A] metadata.google.internal.some-namespace.svc.cluster.local from 10.0.0.83&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;2023-10-10T10:02:36+02:00 I1010 08:02:36.576101   1 nanny.go:146] dnsmasq[4810]: 315108 10.0.0.83/24595 forwarded metadata.google.internal.some-namespace.svc.cluster.local to 127.0.0.1#10053&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;2023-10-10T10:02:36+02:00 I1010 08:02:36.576132   1 nanny.go:146] dnsmasq[4810]: 315108 10.0.0.83/24595 reply metadata.google.internal.some-namespace.svc.cluster.local is NXDOMAIN&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Here we can see exact time (&lt;code&gt;08:02:36.571446&lt;/code&gt;). Which &lt;code&gt;dnsmasq&lt;/code&gt; process logged the line, identified by the process ID (&lt;code&gt;dnsmasq[4811]&lt;/code&gt;). A number identifying an individual request (&lt;code&gt;315207&lt;/code&gt;). And the client IP and port (&lt;code&gt;10.0.0.83/56382&lt;/code&gt;). One process always maps to one TCP connection so PID and client IP and port will always match.&lt;/p&gt;
&lt;p&gt;Just from this snippet we can see that multiple DNS requests are handled by each process and we have exact time for when individual requests were received from the client as well as exact time for replies.&lt;/p&gt;
&lt;p&gt;I wrote a small (and ugly) &lt;a href=&quot;https://github.com/signicat/blog-attachements/blob/main/2023-gke-node-local-dns-cache/files/analyze-dnsmasq-logs/main.go&quot;&gt;program&lt;/a&gt; to analyze the logs. It gathers the duration of each request identified by the request number and each “session” identified by the process ID.&lt;/p&gt;
&lt;p&gt;It filters out requests faster than 2 milliseconds and sessions shorter than 500 milliseconds. The output looks like:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Request: PID 5439 RequestNumber 356558 Domain . Duration 2.273ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Request: PID 4914 RequestNumber 319022 Domain api.statuspage.io.some-namespace.svc.cluster.local Duration 2.619ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Request: PID 4915 RequestNumber 319123 Domain api.statuspage.io.cluster.local Duration 3.088ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Session: PID 4890 Duration 1.005229s&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Session: PID 5022 Duration 852.861ms&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Session: PID 5435 Duration 2.627174s&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The vast majority of individual requests complete in microseconds, and none are slower than 10 milliseconds. This is an indication that the delays aren’t coming from the processing of individual requests.&lt;/p&gt;
&lt;p&gt;Sessions however last much longer, regularly in the 1-3 second range. This isn’t necessarily a problem since it’s resource efficient to keep sessions longer and re-using them for many requests.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;appendix-analyzing-concurrent-tcp-connections&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;appendix---analyzing-concurrent-tcp-connections&quot;&gt;Appendix - Analyzing concurrent TCP connections&lt;/h3&gt;
&lt;p&gt;I want to see how many connections are open at any point (and not impacted by metric scraping intervals etc) as well as how long they tend to stay idle before being closed.&lt;/p&gt;
&lt;p&gt;I made another small (and probably even uglier) &lt;a href=&quot;https://github.com/signicat/blog-attachements/blob/main/2023-gke-node-local-dns-cache/files/analyze-tcp-conns/main.go&quot;&gt;program&lt;/a&gt; to analyze the packet captures (exported from Wireshark as CSV) from &lt;code&gt;kube-dns&lt;/code&gt; and try to answer those questions.&lt;/p&gt;
&lt;p&gt;While iterating through every packet it keeps a counter on how many distinct connections are observed as well as the time since the previous packet in the same connection was observed:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Time: 08:08:30.610010781 TimeShort: 08:08:30 Port: 61236 Action: open Result: none Flags: PSHACK IdleGroup: fast ConIdleTime: 564.26µs ActiveConnections: 11&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Time: 08:08:30.610206151 TimeShort: 08:08:30 Port: 10806 Action: open Result: none Flags: ACK IdleGroup: fast ConIdleTime: 178.4µs ActiveConnections: 11&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Time: 08:08:30.900267009 TimeShort: 08:08:30 Port: 62083 Action: open Result: none Flags: FINACK IdleGroup: slow ConIdleTime: 6.429595796s ActiveConnections: 11&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Plotting &lt;code&gt;ActiveConnections&lt;/code&gt; in the Excel graph we have from &lt;a href=&quot;#analyzing-dns-problems-based-on-packet-capture&quot;&gt;Analyzing DNS problems based on packet captures&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Graph showing active TCP connections in green. DNS resolution time from packet captures in blue. Number of dnsmasq processes in orange&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./dnsmasq-proc-and-tcp-conns-and-packet-latency.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;Up to the limit of 20 &lt;code&gt;dnsmasq&lt;/code&gt; processes they follow the number of open TCP connections pretty closely. However for long periods of time there are way more open connections than &lt;code&gt;dnsmasq&lt;/code&gt; is allowed to spawn new child processes to handle. This also overlaps with the sudden huge increases in latency. Another thing we can infer from this is that new TCP connections are successfully being opened from &lt;code&gt;node-local-dns&lt;/code&gt; (CoreDNS) to &lt;code&gt;dnsmasq&lt;/code&gt;, even though &lt;code&gt;dnsmasq&lt;/code&gt; is unable to handle them yet. Probably the master &lt;code&gt;dnsmasq&lt;/code&gt; process accepts the connections but blocking the request until there is room to spawn a new child process.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;grep&lt;/code&gt;ing for “slow” we get all packets where the connection was idle for more than 1 second:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Time: 08:08:18.064874738 TimeShort: 08:08:18 Port: 19956 Action: open Result: none Flags: FINACK IdleGroup: slow ConIdleTime: 7.535928188s ActiveConnections: 9&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Time: 08:08:18.064888758 TimeShort: 08:08:18 Port: 30168 Action: open Result: none Flags: FINACK IdleGroup: slow ConIdleTime: 7.577724318s ActiveConnections: 9&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Time: 08:08:18.064909758 TimeShort: 08:08:18 Port: 41718 Action: open Result: none Flags: FINACK IdleGroup: slow ConIdleTime: 7.393646609s ActiveConnections: 9&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Time: 08:08:18.064911088 TimeShort: 08:08:18 Port: 30386 Action: open Result: none Flags: FINACK IdleGroup: slow ConIdleTime: 7.535962768s ActiveConnections: 9&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Time: 08:08:18.064920468 TimeShort: 08:08:18 Port: 21482 Action: open Result: none Flags: FINACK IdleGroup: slow ConIdleTime: 7.535946008s ActiveConnections: 9&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This is particularly interesting since &lt;code&gt;node-local-dns&lt;/code&gt; is configured to expire connections after 1 second. And in theory should be cleaned up after at most 2 seconds. I managed to keep myself from diving head first into that particular rabbit hole though.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The underlying &lt;a href=&quot;https://github.com/miekg/dns/&quot;&gt;dns Go library&lt;/a&gt; that CoreDNS uses also has a &lt;a href=&quot;https://github.com/miekg/dns/blob/master/server.go#L18&quot;&gt;limit of 128 DNS queries&lt;/a&gt; for a single TCP connection before closing it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;analyzing-dns-problems-based-on-packet-capture&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;appendix---analyzing-dns-problems-based-on-packet-captures&quot;&gt;Appendix - Analyzing DNS problems based on packet captures&lt;/h3&gt;
&lt;p&gt;Now that we have some packet captures we can start dissecting and analyzing and looking for the needle in the haystack.&lt;/p&gt;
&lt;p&gt;I’ll be using &lt;a href=&quot;https://www.wireshark.org/&quot;&gt;Wireshark&lt;/a&gt; for this.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://www.youtube.com/watch?v=RRjutHjGdCY&quot;&gt;This youtube video&lt;/a&gt; shows a few neat tricks. Thanks!&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;If you are looking at the traffic between &lt;code&gt;dnsmasq&lt;/code&gt; and &lt;code&gt;kube-dns&lt;/code&gt; (&lt;code&gt;lo&lt;/code&gt; network interface) go to &lt;strong&gt;Analyze&lt;/strong&gt; -&amp;gt; &lt;strong&gt;Decode&lt;/strong&gt; as and add a mapping from port 10053 to DNS to have Wireshark decode the packets correctly.&lt;/li&gt;
&lt;li&gt;Start by applying the display filter &lt;code&gt;dns.flags.response == 1&lt;/code&gt; to only show DNS responses.&lt;/li&gt;
&lt;li&gt;Find a DNS Response packet and find the &lt;code&gt;[Time: n.nnn seconds]&lt;/code&gt; field, right click it and &lt;strong&gt;Apply as Column&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;You can also add &lt;code&gt;Transaction ID&lt;/code&gt; and the &lt;code&gt;Name&lt;/code&gt; field (under &lt;strong&gt;Queries&lt;/strong&gt;) as columns as well.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Then you can export as CSV for example in &lt;strong&gt;File&lt;/strong&gt; -&amp;gt; &lt;strong&gt;Export Packet Dissections…&lt;/strong&gt; -&amp;gt; As &lt;strong&gt;CSV&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Importing the data from the &lt;code&gt;node-local-dns&lt;/code&gt; packet capture into Excel (yes, there’s no way escaping Excel!) and plotting the duration of each and every DNS lookup over a 20 minute period, colored by upstream server:&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Graph showing DNS resolution time colored by upstream server based on packet captures on node-local-dns Pod&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./dns-lookup-time-by-upstream-server-over-time.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;We clearly see huge spikes in latency. But it seems to only affect queries being forwarded to the two &lt;code&gt;kube-dns&lt;/code&gt; Pods (&lt;code&gt;172.20.1.38&lt;/code&gt; &amp;amp; &lt;code&gt;172.20.7.57&lt;/code&gt;).&lt;/p&gt;
&lt;p&gt;Another interesting finding is that it appears to happen at the exact same time on both &lt;code&gt;kube-dns&lt;/code&gt; Pods. Weird. If we didn’t already know that (for at least some queries) the added duration happens inside the &lt;code&gt;dnsmasq&lt;/code&gt; container, I would probably suspect a problem on the network.&lt;/p&gt;
&lt;p&gt;Plotting the duration on requests arriving at &lt;code&gt;dnsmasq&lt;/code&gt; on one of the &lt;code&gt;kube-dns&lt;/code&gt; Pods shows the same pattern:&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Graph showing DNS resolution time based on packet captures on kube-dns Pod&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./dnsmasq-latency.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;Another thing I noticed is that sometimes close to when request duration would spike, Wireshark warns about &lt;code&gt;TCP Port numbers reused&lt;/code&gt;:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;42878	19:09:03.987437048	127.0.0.1	127.0.0.1	TCP	78	44516	[TCP Port numbers reused] 44516 → 10053 [SYN] Seq=0 Win=43690 Len=0&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;However Wireshark doesn’t discriminate whether that was 20 minutes or 2 seconds ago. Only that it occurs in the same packet capture.&lt;/p&gt;
&lt;p&gt;One hypothesis I had was that outgoing requests from &lt;code&gt;dnsmasq&lt;/code&gt; to &lt;code&gt;kube-dns&lt;/code&gt; would be stuck waiting for available TCP ports. I plotted the source port usage over time:&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Graph showing TCP source port for connections from node-local-dns based on packet captures&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./dnsmasq-source-port.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;It did not strengthen my suspicion and some random checking shows that the TCP Port reuse is far enough spaced in time (minutes) to avoid problems. So for the time being I’m not digging further into this.&lt;/p&gt;
&lt;h2 id=&quot;footnotes&quot;&gt;Footnotes&lt;/h2&gt;
&lt;p&gt;&lt;a id=&quot;footnote-a&quot;&gt;&lt;/a&gt;
&lt;strong&gt;[A]&lt;/strong&gt; This doesn’t strictly mean that 20 new TCP connections are required since many of them are probably retries.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;footnote-b&quot;&gt;&lt;/a&gt;
&lt;strong&gt;[B]&lt;/strong&gt; Note that even if the graph doesn’t hit 22 you may still be affected. The number of processes is counted only at the exact time the metrics are scraped, in our case 10 seconds. You can manually sample the process count by attaching a debug container to &lt;code&gt;dnsmasq&lt;/code&gt; (&lt;code&gt;kubectl debug -c debug -n kube-system -it kube-dns-6fb7c8866c-bxj7f --image=ubuntu:22.04 --target=dnsmasq&lt;/code&gt;) and running &lt;code&gt;for i in $(seq 1 1800) ; do echo &quot;$(date) Try: ${i} DnsmasqProcess: $(pidof dnsmasq | wc -w)&quot;; sleep 1; done&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;footnote-c&quot;&gt;&lt;/a&gt;
&lt;strong&gt;[C]&lt;/strong&gt; CPU throttling would not be an issue in this case since &lt;code&gt;kube-dns&lt;/code&gt; does not have CPU limits set.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;footnote-d&quot;&gt;&lt;/a&gt;
&lt;strong&gt;[D]&lt;/strong&gt; Although you should be cautious of calling services by name in other namespaces if they are owned by different teams. As that introduces coupling on what should be implementation details across team boundaries. In Signicat we always call the full FQDN and path if connecting to services owned by other teams.&lt;/p&gt;
</content:encoded><category>tech</category><category>kubernetes</category><category>gke</category><category>dns</category></item><item><title>Kubernetes Sidecar Config Drift</title><link>https://blog.stian.omg.lol/p/kubernetes-sidecar-config-drift/</link><guid isPermaLink="true">https://blog.stian.omg.lol/p/kubernetes-sidecar-config-drift/</guid><description>When using a Sidecar Injector (such as Istio), there is nothing that ensures that an update (potentially breaking) to a sidecar config template is applied/updated on Pods that have already been injected with a sidecar. This post describes the causes of this problem, as well as introducing a tool to mitigate it.</description><pubDate>Tue, 03 May 2022 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://blog.stian.omg.lol/_astro/2022-05-03-kubernetes-sidecar-config-drift.DqqvBfzL_ZFI2Xz.jpg&quot; width=&quot;1600&quot; height=&quot;914&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;When using a Sidecar Injector (such as Istio), there is nothing that ensures that an update (potentially breaking) to a sidecar config template is applied/updated on Pods that have already been injected with a sidecar. This post describes the causes of this problem, as well as introducing a tool to mitigate it.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;This is a cross-post of a blog post also published on the &lt;a href=&quot;https://www.signicat.com/blog/kubernetes-your-sidecar-configurations-are-drifting&quot;&gt;Signicat Blog&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Huge thanks to one of my favorite clients, &lt;a href=&quot;https://www.signicat.com/&quot;&gt;Signicat&lt;/a&gt;, and especially &lt;a href=&quot;https://www.linkedin.com/in/jon-skarpeteig/&quot;&gt;Jon&lt;/a&gt;, for allowing me to share some of the nitty gritty details of a challenge that I believe is probably quite widespread, yet under-appreciated, in modern Kubernetes cloud environments.&lt;/p&gt;
&lt;p&gt;Last week, working on Signicat’s next generation cloud platform, I discovered that several individuals invented their own ways of mitigating what I now call Sidecar Configuration Drift.&lt;/p&gt;
&lt;p&gt;To ease the pain I created &lt;a href=&quot;https://github.com/StianOvrevage/k8s-sidecar-rollout&quot;&gt;k8s-sidecar-rollout&lt;/a&gt; to restart the required workloads and thereby updating their sidecar configurations.&lt;/p&gt;
&lt;p&gt;This blog post is a bit of background information on what the causes of this problem is.&lt;/p&gt;
&lt;h2 id=&quot;what-is-sidecar-config-drift&quot;&gt;What is Sidecar Config Drift?&lt;/h2&gt;
&lt;p&gt;When using a Sidecar Injector (such as Istio), there is nothing that ensures that an update (potentially breaking) to a sidecar config template is applied/updated on Pods that have already been injected with a sidecar.&lt;/p&gt;
&lt;p&gt;This means that after updating a sidecar config it may take a very long time until all Pods have the updated config. They may receive the updates at any undetermined time in the future. While the updates are pending things might not work as expected and things may not be compliant. When the update is finally applied to the Pod it may surface breaking changes.&lt;/p&gt;
&lt;p&gt;I call this phenomena &lt;em&gt;Sidecar Configuration Drift&lt;/em&gt;.&lt;/p&gt;
&lt;h2 id=&quot;background&quot;&gt;Background&lt;/h2&gt;
&lt;h3 id=&quot;kubernetes---declarative-vs-imperative&quot;&gt;Kubernetes - Declarative vs imperative&lt;/h3&gt;
&lt;p&gt;What makes Kubernetes so powerful is also what can make it hard and confusing to work with until your mindset has shifted.&lt;/p&gt;
&lt;p&gt;That is the philosophy of being declarative instead of imperative like most of us are used to for the past 20 years.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Declarative means you tell Kubernetes HOW you want things to look. You DON’T tell Kubernetes WHAT to do.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For example you do not tell Kubernetes to scale your Deployment to 5 instances. You tell Kubernetes that your Deployment should have 5 instances. The difference is subtle but extremely important. In a highly dynamic cloud environment your instances might disappear or crash for a multitude of reasons. In this declarative mindset you should not care about that since Kubernetes is tasked with ensuring 5 instances and will (try to) provision new ones when it’s needed.&lt;/p&gt;
&lt;p&gt;This is the apparent magic which makes Kubernetes.&lt;/p&gt;
&lt;p&gt;This magic is technically solved with what we call &lt;strong&gt;Controllers&lt;/strong&gt;. A controller is responsible for constantly comparing the &lt;strong&gt;Actual state&lt;/strong&gt; and &lt;strong&gt;Desired state&lt;/strong&gt;. To scale your Deployment to 5 instances you set &lt;code&gt;replicas: 5&lt;/code&gt; on the &lt;code&gt;Deployment&lt;/code&gt; resource. If needed a controller will create a completely new &lt;code&gt;ReplicaSet&lt;/code&gt; resource with 5 instances. Another controller will then create 5 &lt;code&gt;Pods&lt;/code&gt;. And the scheduler will finally try to place and start those &lt;code&gt;Pods&lt;/code&gt; on actual nodes. The &lt;code&gt;ReplicaSet&lt;/code&gt; controller will scale down the old once the new has reached it’s desired state.&lt;/p&gt;
&lt;p&gt;This constant process is called a &lt;strong&gt;reconciliation loop&lt;/strong&gt; and it’s a critical feature.&lt;/p&gt;
&lt;h3 id=&quot;sidecars&quot;&gt;Sidecars&lt;/h3&gt;
&lt;p&gt;Sidecar is a Kubernetes design pattern. A sidecar is simply a container that is living side-by-side with the main application container that can do tasks that are logically not part of the application. Examples of this can be log handling, monitoring agents, proxies. &lt;strong&gt;All containers in a Pod (including sidecars) share the same filesystem, kernel namespace, IP addresses etc.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;More about sidecars: &lt;a href=&quot;https://kubernetes.io/blog/2015/06/the-distributed-system-toolkit-patterns/&quot;&gt;The Distributed System ToolKit: Patterns for Composite Containers&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;sidecar-injection&quot;&gt;Sidecar injection&lt;/h3&gt;
&lt;p&gt;Sidecar injection is when a sidecar container is added to a Pod &lt;strong&gt;even if the sidecar isn’t defined in any of the higher level primitives, such as &lt;code&gt;Deployment&lt;/code&gt; or &lt;code&gt;StatefulSet&lt;/code&gt;.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;When for example the &lt;code&gt;ReplicaSet&lt;/code&gt; controller will &lt;code&gt;Create&lt;/code&gt; a new &lt;code&gt;Pod&lt;/code&gt;. If configured, the Kubernetes API will call one or more &lt;code&gt;MutatingWebhooks&lt;/code&gt;. These webhooks can then change the &lt;code&gt;Pod&lt;/code&gt; definition before they are saved (and picked up by the scheduler).&lt;/p&gt;
&lt;h3 id=&quot;the-problem&quot;&gt;The Problem&lt;/h3&gt;
&lt;p&gt;The problem is when updating a sidecar injection template there is no system that runs a reconciliation loop.&lt;/p&gt;
&lt;p&gt;The webhook just updates the &lt;code&gt;Pod&lt;/code&gt; template. It does not keep track of which &lt;code&gt;Pods&lt;/code&gt; have gotten which template or check if any template change would result in a different sidecar configuration.&lt;/p&gt;
&lt;p&gt;The controllers ALSO does not continuously monitor if and how re-creating the same &lt;code&gt;Pod&lt;/code&gt; (without sidecars) would result in a different Pod once the sidecars have been injected.&lt;/p&gt;
&lt;p&gt;In effect sidecar injection does not follow the expected declarative pattern that the rest of Kubernetes does.&lt;/p&gt;
&lt;h3 id=&quot;consequences&quot;&gt;Consequences&lt;/h3&gt;
&lt;p&gt;If a platform team changes the istio sidecar template it will not actually take effect on a &lt;code&gt;Pod&lt;/code&gt; until that &lt;code&gt;Pod&lt;/code&gt; for some reason is re-created.&lt;/p&gt;
&lt;p&gt;Let’s assume the &lt;code&gt;istio-proxy&lt;/code&gt; sidecar template have been updated by the platform team. We roll it out and test it and it seems to work. But the change will break some applications running in the cluster.&lt;/p&gt;
&lt;p&gt;That breakage will go un-noticed until:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;The product team commits changes that triggers a re-deploy. The deployment will suddenly fail but it might not have anything to do with the actual changes the team did to the application. This is surely confusing!&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;The platform team for example upgrades a pool of worker nodes causing all `Pods` to be re-created on new nodes.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;A Pod is re-created when a Kubernetes worker node crashes. In this scenario it appears the failure spawned into existence out of nowhere since neither the product team nor platform team actually “did” anything to trigger it.&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;Also worth noting is that any attempts at Rolling back a Deployment now containing failing Pods will not actually fix anything.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;It’s the sidecar templates that needs to be rolled back and Pods probably need to be re-created again.&lt;/p&gt;
&lt;h3 id=&quot;mitigations&quot;&gt;Mitigations&lt;/h3&gt;
&lt;p&gt;We can mitigate drift by:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Re-starting all Pods in the cluster whenever we update sidecar injection templates.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Sometimes we might forget to re-start so regularly re-start all Pods in the cluster anyway.&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;k8s-sidecar-rollout&quot;&gt;k8s-sidecar-rollout&lt;/h2&gt;
&lt;p&gt;To make these restarts easy and fast I’ve created &lt;a href=&quot;https://github.com/StianOvrevage/k8s-sidecar-rollout&quot;&gt;https://github.com/StianOvrevage/k8s-sidecar-rollout&lt;/a&gt; .&lt;/p&gt;
&lt;p&gt;It’s a tool that figures out (with your help) which workloads (Deployment, StatefulSet, DaemonSet) that needs to be rolled out again (re-started) and then rolls out for you. Head over to the GitHub repo for installation and complete usage instructions. Here is an example of how it can be used:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;python3 sidecar-rollout.py \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --sidecar-container-name=istio-proxy \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --include-daemonset=true \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --annotation-prefix=myCompany \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --parallel-rollouts 10 \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --only-started-before=&quot;2022-05-01 13:00&quot; \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --exclude-namespace=kube-system \&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    --confirm=true&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This will gather all Pods with a container named &lt;code&gt;istio-sidecar&lt;/code&gt; belonging to a Deployment or DaemonSet that was started before 2022-05-01 13:00 (which may be when we updated the istio sidecar config template) excluding the &lt;code&gt;kube-system&lt;/code&gt; namespace. It will patch the workloads with two annotations with &lt;code&gt;myCompany&lt;/code&gt; prefix and run 10 rollouts in parallel.&lt;/p&gt;
&lt;p&gt;The script that now re-starts Pods adds two annotations indicating that a restart to update sidecars has occurred as well as the time:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;$ kubectl get pods -n product-team some-api-7cdc65482b-ged13 -o yaml | yq &apos;.metadata.annotations&apos;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    sidecarRollout.rollout.timestamp: 2022-05-03T18:05:31&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    sidecarRollout.rollout.reason: Update sidecars istio-proxy&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The idea is that if your Pods are suddenly failing, you can quickly check the annotations and see if it has anything to do with sidecar updates or not.&lt;/p&gt;
&lt;p&gt;These annotations will of course disappear again when a Deployment is updated.&lt;/p&gt;
</content:encoded><category>tech</category><category>kubernetes</category><category>istio</category></item><item><title>Yak shaving - Photo drips for my mom</title><link>https://blog.stian.omg.lol/p/yak-shaving-photo-drips-for-my-mom/</link><guid isPermaLink="true">https://blog.stian.omg.lol/p/yak-shaving-photo-drips-for-my-mom/</guid><description>TL;DR: I finally organized my photo archive and in an evening created a service to e-mail my mom a photo from the last 20 years every morning.</description><pubDate>Sun, 13 Feb 2022 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://blog.stian.omg.lol/_astro/2022-02-13-photo-drips-for-my-mom.DGDIiuxJ_Z19OCr7.jpg&quot; width=&quot;1600&quot; height=&quot;914&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;TL;DR: I finally organized my photo archive and in an evening created a service to e-mail my mom a photo from the last 20 years every morning.&lt;/p&gt;
&lt;h2 id=&quot;yak-shaving---photo-drips-for-my-mom&quot;&gt;Yak shaving - Photo drips for my mom&lt;/h2&gt;
&lt;p&gt;Update: Check out &lt;a href=&quot;https://github.com/StianOvrevage/photo-drips&quot;&gt;https://github.com/StianOvrevage/photo-drips&lt;/a&gt; for ugly but working code.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Update May 2022: My mom told me this week that I must NEVER stop sending these daily photos &amp;lt;3&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3 id=&quot;background&quot;&gt;Background&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;TL;DR: I finally organized my photo archive and in an evening created a service to e-mail my mom a photo from the last 20 years every morning.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This started out a few weeks ago when I decided to reinstall Windows on my laptop.&lt;/p&gt;
&lt;p&gt;The first hurdle was that my 2-3-4 different cloud storage subscriptions were all overdue for a clean-up. The one best suited had me throttled to 1Mbit/s since I was storing 20TB+ of data.&lt;/p&gt;
&lt;p&gt;I’ve been taking a lot of photos and videos since I got my first digital camera 22 years ago. Even though I use Lightroom to keep some order there was some duplication and discontinuity. So (after upgrading my fibreoptic internet to 1Gbit, optimizing my home network and cleaning up space on my machine) I started to clean up and organize the various catalogues.&lt;/p&gt;
&lt;p&gt;Looking at memories from 20 years ago made me realize I’m better at taking photos than “utilizing” them afterwards. What really is the point, then? I took a picture of one on the screen with my phone and sent on snapchat to my mom, not thinking much about it. But she was really thrilled, which made me really happy as well.&lt;/p&gt;
&lt;p&gt;For my current client I’m nearing the end of my contract and for the last two months I’ve mainly been maintaining, documenting and handing over and I really miss building things and solving problems.&lt;/p&gt;
&lt;p&gt;Recently a potential client asked about Python and AWS Lambda competence. It’s not something I work with daily and it’s not on my CV. But it made me think about all the various languages and tools I’ve used during the last decade.&lt;/p&gt;
&lt;p&gt;My subconscious brain offers up an idea on how to a) build something b) brush up some Python and Lambda knowledge and c) brighten my moms day.&lt;/p&gt;
&lt;h4 id=&quot;concept&quot;&gt;Concept&lt;/h4&gt;
&lt;p&gt;Every morning a photo from my archives is e-mailed to my mom, a photo drip.&lt;/p&gt;
&lt;p&gt;There is also a gallery where she can look at the previous photo drips.&lt;/p&gt;
&lt;h4 id=&quot;process-and-goals&quot;&gt;Process and goals&lt;/h4&gt;
&lt;p&gt;The primary objective was to have a working prototype as quickly as possible. So no premature optimization, refactoring or anything.&lt;/p&gt;
&lt;h3 id=&quot;photo-selection-and-preparation&quot;&gt;Photo selection and preparation&lt;/h3&gt;
&lt;p&gt;At around 3pm I started browsing photos from year 2000. It took me about 75 minutes to pick out about 300 photos from the first 20.000.&lt;/p&gt;
&lt;p&gt;Tip: When browsing in Lightroom, press B to add to Quick Collection.&lt;/p&gt;
&lt;p&gt;Export the pictures in the quick collection with a custom filename format like this &lt;code&gt;0003-2000-05-14.jpg&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;This format ensures filenames are ordered from oldest to newest, and the date can later be extracted directly from the filename without doing any EXIF stuff.&lt;/p&gt;
&lt;p&gt;Photos resized to 1500x1500. Never enlarge. Sharpen for screen.&lt;/p&gt;
&lt;h3 id=&quot;photo-storage&quot;&gt;Photo storage&lt;/h3&gt;
&lt;p&gt;It would have been quicker to set up a AWS EC2 virtual machine running all the components but that would be too easy.&lt;/p&gt;
&lt;p&gt;The requirements for storage is: cheap, reliable, publicly available. So a standard AWS S3 bucket should do just fine.&lt;/p&gt;
&lt;p&gt;Using the browser I upload all the photos in batch to a new bucket. In a sub-folder with a random name. That should provide an appropriate level of security and avoid strangers on the internet stumbling upon it. Since the bucket is publicly available.&lt;/p&gt;
&lt;p&gt;I verify that I can load the pictures in my browser with the public S3 URL. It took a few attempts at getting the permissions right, and I suspect adding this policy was required even though everything in the settings was set to “Full public access”.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;{&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &quot;Version&quot;: &quot;2012-10-17&quot;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &quot;Statement&quot;: [&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;            &quot;Sid&quot;: &quot;PublicRead&quot;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;            &quot;Effect&quot;: &quot;Allow&quot;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;            &quot;Principal&quot;: &quot;*&quot;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;            &quot;Action&quot;: [&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;                &quot;s3:GetObject&quot;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;                &quot;s3:GetObjectVersion&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;            ],&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;            &quot;Resource&quot;: [&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;                &quot;arn:aws:s3:::memorydrops/*&quot;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;            ]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        }&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    ]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;sending-the-e-mail&quot;&gt;Sending the e-mail&lt;/h3&gt;
&lt;p&gt;Sending e-mails directly with SMTP is really not an option anymore because of all the spam blocking systems. I know we need to use a third party service or something.&lt;/p&gt;
&lt;p&gt;I first looked at AWS SES (simple-email-service). But nothing about it seemed simple. Then looked at a few of the stablished ones such as SendGrid and Mailgun. These tools have evolved a lot and now has a plethora of features aimed at marketing, transactional e-mails etc. Probably overkill and potentially time consuming while having no transferable knowledge or code when hitting a dead-end since they are all different APIs.&lt;/p&gt;
&lt;p&gt;What about just using my own Gmail SMTP to ship?&lt;/p&gt;
&lt;p&gt;The first tutorials about that required turning on “Less secure app access” for my account.&lt;/p&gt;
&lt;p&gt;Not accepting that trade-off, I found &lt;a href=&quot;https://levelup.gitconnected.com/an-alternative-way-to-send-emails-in-python-5630a7efbe84&quot;&gt;https://levelup.gitconnected.com/an-alternative-way-to-send-emails-in-python-5630a7efbe84&lt;/a&gt; where I learned that I can generate an “App password”, similar to an API Key for specific Google applications without disabling other security features. Jackpot!&lt;/p&gt;
&lt;p&gt;The article also has working code that I shamelessly used as a starting point.&lt;/p&gt;
&lt;p&gt;First I attached the photo of the day, but I really want it embedded. Even though I don’t really like linking images on a public URL (S3 bucket) because of potential browser and client issues that’s what I ended up with. The alternative of base64 encoding the data inline seemed like a chore. Besides I know mom uses Gmail in Chrome anyway so if it worked for me it would probably work for her.&lt;/p&gt;
&lt;p&gt;I don’t want the system to be dependent on any state management, databases, etc.&lt;/p&gt;
&lt;p&gt;Picking the right photo each day is simply select file N, where N is the days since the first deployment.&lt;/p&gt;
&lt;p&gt;Ideally I would use AWS S3 API to list the contents of the bucket to get the available files. But to save some time the “photo index” is simply a text file listing the files:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;ls &amp;gt; filelist.txt&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;I create a new Lambda function and paste the python script and filelist.txt directly in the Lambda code editor and deploy and test it. Working!&lt;/p&gt;
&lt;p&gt;The Lambda function is the configured with a EventBridge trigger with the schedule &lt;code&gt;cron(0 5 * * ? *)&lt;/code&gt; that should trigger the function at 0500 UTC every day.&lt;/p&gt;
&lt;h3 id=&quot;gallery&quot;&gt;Gallery&lt;/h3&gt;
&lt;p&gt;I also want to have a gallery where she can view the previous memory drops without having to shuffle through endless emails.&lt;/p&gt;
&lt;p&gt;Started out by looking at VueJS for the frontend, which I have used before. But I have not used the new version yet and I suspected it might take a long time getting a project set up from scratch since unfortunately tutorials, tips, documentation in the frontend / JavaScript world has a tendency to be chronically outdated and unreliable.&lt;/p&gt;
&lt;p&gt;Dropped that and opted for a very simple native JS photo gallery called &lt;a href=&quot;https://github.com/ericleong/zoomwall.js/&quot;&gt;zoomwall.js&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Only showing a selection of photos depending on the day however requires a bit more engineering. (Writing this I realize another acceptable approach would be to use native JS to manipulate the DOM directly.)&lt;/p&gt;
&lt;p&gt;I implemented this by inlining the CSS and JS into one template HTML file that is rendered on-demand in a Python Lambda function. Then using AWS API Gateway to expose it as a normal webserver.&lt;/p&gt;
&lt;p&gt;I wanted to use Jinja2 for templating instead of the built-in Python templating functions. Doing that causes some headache since the Lambda environment does not have Jinja2 installed.&lt;/p&gt;
&lt;p&gt;Luckily creating a custom deployment package (.zip) including dependencies is rather trivial.&lt;/p&gt;
&lt;h3 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h3&gt;
&lt;p&gt;All in all this project was started with filtering photos at 3pm, having dinner from 5pm to 6pm and by 8.30pm everything was deployed. The following morning I had the memory drop in my inbox to great delight. An unexpected benefit of using my own Gmail for sending the e-mail is that mom can reply directly.&lt;/p&gt;
&lt;p&gt;Expected costs&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;S3 offers 5GB of standard storage for free for 12 months. After that I expect the cost to be in the $0.x range per month.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;API Gateway offers 1 Million API calls per month for free for 12 months. After that I expect the cost to be in the $0.x range per month.&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Lambda offers 1 Million requests per month forever.&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Performance has been sacrificed for the gallery to make it simple. Server-side rendering is not going to be as fast as client-side. Python is not the fastest alternative. Using a FaaS such as Lambda also introduces penalties and unknowns (cold starts, etc). Yet the gallery HTML loads in ~200ms and and seems instant.&lt;/p&gt;
&lt;h3 id=&quot;improvements&quot;&gt;Improvements&lt;/h3&gt;
&lt;p&gt;After deciding to share the code I spent an hour or two writing this document as well as some necessary clean-up and changes from the prototype.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;filelist.txt&lt;/code&gt; is not bundled with the code. It’s hosted in the S3 bucket along with the photos. That means I can update and add photos later without touching code.&lt;/p&gt;
&lt;p&gt;That requires the &lt;code&gt;requests&lt;/code&gt; package, so now both modules are packaged with dependencies and uploaded via AWS CLI instead of browser.&lt;/p&gt;
&lt;p&gt;Some hard-coded URLs etc have been converted to environment variables.&lt;/p&gt;
&lt;p&gt;Added the &lt;code&gt;exif&lt;/code&gt; Python package to extract the original time and date of the photo to include in the e-mail.&lt;/p&gt;
&lt;p&gt;Added an URL redirect from a prettier domain to the auto generated API Gateway hostname of the gallery. Did not bother with proper custom domain since that requires a lot of fiddling with SSL certificates.&lt;/p&gt;
&lt;h3 id=&quot;bugs-and-future-improvements&quot;&gt;Bugs and future improvements&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;The JS gallery seems slightly broken. Might replace with a better one.&lt;/li&gt;
&lt;li&gt;The e-mail HTML is not pretty.&lt;/li&gt;
&lt;li&gt;Make sender and recipient e-mails environment variables (but recipient is currently a Python list, so).&lt;/li&gt;
&lt;li&gt;Make start date env var. Requires parsing date from user.&lt;/li&gt;
&lt;li&gt;The usual: Check that required env vars are set on startup. Improve error handling and logging.&lt;/li&gt;
&lt;/ul&gt;
</content:encoded><category>tech</category><category>side-quests</category><category>photography</category></item><item><title>A side quest in API development, observability, Kubernetes and cloud with a hint of database</title><link>https://blog.stian.omg.lol/p/a-side-quest-in-api-development-observability-kubernetes-and-cloud-with-a-hint-of-database/</link><guid isPermaLink="true">https://blog.stian.omg.lol/p/a-side-quest-in-api-development-observability-kubernetes-and-cloud-with-a-hint-of-database/</guid><description>Quite often people ask me what I actually do. I have a hard time giving a short answer. Even to colleagues and friends in the industry. Here I will try to show and tell how I spent an evening digging around in a system I helped build for a client.</description><pubDate>Sat, 06 Mar 2021 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://blog.stian.omg.lol/_astro/2021-03-06-a-side-quest-in-api-dev-operations-cloud-and-database.BJvjEF0d_ZGYqBS.jpg&quot; width=&quot;1600&quot; height=&quot;914&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;Quite often people ask me what I actually do. I have a hard time giving a short answer. Even to colleagues and friends in the industry. Here I will try to show and tell how I spent an evening digging around in a system I helped build for a client.&lt;/p&gt;
&lt;p&gt;Quite often people ask me what I actually do. I have a hard time giving a short answer. Even to colleagues and friends in the industry.&lt;/p&gt;
&lt;p&gt;Here I will try to show and tell how I spent an evening digging around in a system I helped build for a client.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;Table of contents&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#Background&quot;&gt;Background&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#TheProblem&quot;&gt;The (initial) problem&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#FixingTheProblem&quot;&gt;Fixing the (initial) problem&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#VerifyingTheFix&quot;&gt;Verifying the (initial) fix&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#BaselineSimple&quot;&gt;Baseline simple request - HTTP1 1 connections, 20000 requests&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#BaselineComplex&quot;&gt;Baseline complex request - HTTP1 1 connections, 20000 requests&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#VerifyingWorkload&quot;&gt;Verifying the fix for assumed workload&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#ComplexHttp1&quot;&gt;Complex request - HTTP1 6 connections, 500 requests&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#ComplexHttp2&quot;&gt;Complex request - HTTP2 500 “connections”, 500 requests&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#DatabaseOptimizations&quot;&gt;Side quest: Database optimizations&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#NextBottleneck&quot;&gt;Determining the next bottleneck&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#ClusterResources&quot;&gt;Side quest: Cluster resources and burstable VMs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Conclusion&quot;&gt;Conclusion&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a id=&quot;Background&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;background&quot;&gt;Background&lt;/h2&gt;
&lt;p&gt;I’m a consultant doing development, DevOps and cloud infrastructure.&lt;/p&gt;
&lt;p&gt;For this specific client I mainly develop APIs using Golang to support new products and features as well as various exporting, importing and processing of data in the background.&lt;/p&gt;
&lt;p&gt;I’m also the “ops” guy handling everything in AWS, setting up and maintaing databases, making sure the “DevOps” works and the frontend and analytics people can do their work with little friction.
99% of the time things work just fine. No data is lost. The systems very rarely have unforeseen downtime and the users can access the data they want with acceptable latency rarely exceeding 500ms.&lt;/p&gt;
&lt;p&gt;A couple of times a year I assess the status of the architecture and set up new environments from scratch and update any documentation that has drifted. This is also a good time to do changes and add or remove constraints in anticipation of future business needs.&lt;/p&gt;
&lt;p&gt;In short, the current tech stack that has evolved over a couple of years is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Everything hosted on Amazon Web Services (AWS).&lt;/li&gt;
&lt;li&gt;AWS managed Elastic Kubernetes Service (EKS) currently on K8s 1.18.&lt;/li&gt;
&lt;li&gt;GitHub Actions for building Docker images for frontends, backends and other systems.&lt;/li&gt;
&lt;li&gt;AWS Elastic Container Registry for storing Docker images.&lt;/li&gt;
&lt;li&gt;Deployment of each system defined as a Helm chart alongside source code.&lt;/li&gt;
&lt;li&gt;Actual environment configuration (Helm values) stored in repo along source code. Updated by GitHub Actions.&lt;/li&gt;
&lt;li&gt;ArgoCD in cluster to manage status of all environments and deployments. Development environments usually automatically deployed on change. Push a button to deploy to Production.&lt;/li&gt;
&lt;li&gt;Prometheus for storing metrics from the cluster and nodes itself as well as custom metrics for our own systems.&lt;/li&gt;
&lt;li&gt;Loki for storing logs. Makes it easier to retrieve logs from past Pods and aggregate across multiple Pods.&lt;/li&gt;
&lt;li&gt;Elastic APM server for tracing.&lt;/li&gt;
&lt;li&gt;Pyroscope for live CPU profiling/tracing of Go applications.&lt;/li&gt;
&lt;li&gt;Betteruptime.com for tracking uptime and hosting status pages.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I might write up a longer post about the details if anyone is interested.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;TheProblem&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;the-initial-problem&quot;&gt;The (initial) problem&lt;/h2&gt;
&lt;p&gt;A week ago I upgraded our API from version 1, that was deployed in January, to version 2 with new features and better architecture.&lt;/p&gt;
&lt;p&gt;One of the endpoints of the API returns an analysis of an object we track. I have previously reduced the amount of database queries by 90% but it still requires about 50 database calls from three different databases.
Getting and analyzing the data usually completes in about 3-400 milliseconds returning an 11.000 line JSON.&lt;/p&gt;
&lt;p&gt;It’s also possible to just call &lt;code&gt;/objects/analysis&lt;/code&gt; to get the analysis for all the 500 objects we are tracking. It takes 20 seconds but is meant for exports to other processes and not interactive use, so not a problem.&lt;/p&gt;
&lt;p&gt;Since the product is under very active development the frontend guys just download the whole analysis for an object to show certain relevant information to users. It’s too early to decide on which information is needed more often and how to optimize for that. Not a problem.&lt;/p&gt;
&lt;p&gt;So we need an overview of some fields from multiple objects in a dashboard / list. We can easily pull analysis from 20 objects without any noticable delay.&lt;/p&gt;
&lt;p&gt;But what if we just want to show more, 50? 200? 500? The frontend already have the IDs for all the objects and fetches them from &lt;code&gt;/objects/id/analysis&lt;/code&gt;. So they loop over the IDs and fire of requests simultaneously.&lt;/p&gt;
&lt;p&gt;Analyzing the network waterfall in Chrome DevTools indicated that the requests now took 20-30 seconds to complete! But looking closer most of the time they were actually queued up in the browser. This is because
Chrome only allows 6 concurrent TCP connection to the same origin when using HTTP1 (&lt;a href=&quot;https://developers.google.com/web/tools/chrome-devtools/network/understanding-resource-timing&quot;&gt;https://developers.google.com/web/tools/chrome-devtools/network/understanding-resource-timing&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;FixingTheProblem&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;fixing-the-initial-problem&quot;&gt;Fixing the (initial) problem&lt;/h3&gt;
&lt;p&gt;HTTP2 should fix this problem easily. By default HTTP2 is disabled in nginx-ingress. I add a couple of lines enabling it and update the Helm deployment of the ingress controller.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;VerifyingTheFix&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;verifying-the-initial-fix&quot;&gt;Verifying the (initial) fix&lt;/h3&gt;
&lt;p&gt;Some common development tools doesn’t support HTTP2, such as Postman. So I found &lt;code&gt;h2load&lt;/code&gt; which can both help me verify HTTP2 is working and I also get to measure the improvement, nice!&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Note that I’m not using the analysis endpoint since I want to measure the change from HTTP1 to HTTP2 and it will become apparent later that there are other bottlenecks preventing us from a linear performance increase when just changing from HTTP1 to HTTP2.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;Also note that this is somewhat naive since it requests the same URL over and over which can give false results due to any caching. But fortunately we don’t do any caching yet.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;BaselineSimple&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;baseline-simple-request---http1-1-connections-20000-requests&quot;&gt;Baseline simple request - HTTP1 1 connections, 20000 requests&lt;/h4&gt;
&lt;p&gt;Using 1 concurrent streams, 1 client and HTTP1 I get an estimate of performance pre-http2:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;h2load --h1 --requests=20000 --clients=1 --max-concurrent-streams=1 https://api.x.com/api/v1/objects/1&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The results are as expected:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;finished in 1138.99s, 17.56 req/s, 18.41KB/s&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;requests: 20000 total, 20000 started, 20000 done, 19995 succeeded, 5 failed, 0 errored, 0 timeout&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Overview from Elastic APM. Duration is very acceptable at around 20ms. No errors. And about 25% of the time spent doing database queries.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./0-baseline-http1-1-concurrent-apm.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Container CPU usage. Nothing special.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./0-baseline-http1-1-concurrent-cpu.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Database query latency. The vast majority under 5ms. Acceptable.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./0-baseline-http1-1-concurrent-db-latency.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Number of DB queries per second.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./0-baseline-http1-1-concurrent-db-queries.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;HTTP response latency.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./0-baseline-http1-1-concurrent-http-latency.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Number of HTTP requests per second. Unsurprisingly the number of database queries are identical to the number of HTTP requests. Latency of HTTP requests also tracks the latency of the (single) database query.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./0-baseline-http1-1-concurrent-http-requests.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;For http2 we set max concurrent streams to the same as number of requests:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;h2load --requests=200 --clients=1 --max-concurrent-streams=200 https://api.x.com/api/v1/objects/1&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Which results in almost half the latency:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;finished in 1.23s, 162.65 req/s, 158.06KB/s&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;requests: 200 total, 200 started, 200 done, 200 succeeded, 0 failed, 0 errored, 0 timeout&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;So HTTP2 is working and providing significant latency improvements. Success!&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;BaselineComplex&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;baseline-complex-request---http1-1-connections-20000-requests&quot;&gt;Baseline complex request - HTTP1 1 connections, 20000 requests&lt;/h4&gt;
&lt;p&gt;We start by establishing a baseline with 1 connection querying over and over.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;h2load --h1 --requests=20000 --clients=1 --max-concurrent-streams=1&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Latency increases as much more computation is done and data is returned. But the latency is consistent which is good. We also see that the database is becomming the bottleneck for where most time is spent.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./1-baseline-http1-1-concurrent-analysis-apm.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;CPU usage increased to 15%. Lower increase than expected considering the complexity involved in serving the requests.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./1-baseline-http1-1-concurrent-analysis-cpu.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Database query latency still mostly under 5ms.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./1-baseline-http1-1-concurrent-analysis-db-latency.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Number of database queries increases by a factor of 10 compared to HTTP requests.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./1-baseline-http1-1-concurrent-analysis-db-queries.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;HTTP latency.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./1-baseline-http1-1-concurrent-analysis-http-latency.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;HTTP requests per second.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./1-baseline-http1-1-concurrent-analysis-http-requests.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;VerifyingWorkload&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;verifying-the-fix-for-assumed-workload&quot;&gt;Verifying the fix for assumed workload&lt;/h3&gt;
&lt;p&gt;So we verified that HTTP2 gives us a performance boost. But what happens when we fire away 500 requests to the much heavier &lt;code&gt;/analysis&lt;/code&gt; endpoint?&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;These graphs are not as pretty since the ones above. This is mainly due to the sampling interval of the metrics and that we need several datapoints to accurately determine the rate() of a counter.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;ComplexHttp1&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;complex-request---http1-6-connections-500-requests&quot;&gt;Complex request - HTTP1 6 connections, 500 requests&lt;/h4&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;finished in 32.25s, 14.88 req/s, 2.29MB/s&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;requests: 500 total, 500 started, 500 done, 500 succeeded, 0 failed, 0 errored, 0 timeout&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;2-burst-http1-6-concurrent-analysis-apm&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./2-burst-http1-6-concurrent-analysis-apm.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;2-burst-http1-6-concurrent-analysis-cpu&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./2-burst-http1-6-concurrent-analysis-cpu.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;2-burst-http1-6-concurrent-analysis-db-latency&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./2-burst-http1-6-concurrent-analysis-db-latency.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;2-burst-http1-6-concurrent-analysis-db-queries&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./2-burst-http1-6-concurrent-analysis-db-queries.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;2-burst-http1-6-concurrent-analysis-http-latency&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./2-burst-http1-6-concurrent-analysis-http-latency.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;2-burst-http1-6-concurrent-analysis-http-requests&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./2-burst-http1-6-concurrent-analysis-http-requests.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;In summary it so far seems to scale linearly with load. Most of the time is spent fetching data from the database. Still very predictable low latency on database queries and the resulting HTTP response.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;ComplexHttp2&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;complex-request---http2-500-connections-500-requests&quot;&gt;Complex request - HTTP2 500 “connections”, 500 requests&lt;/h4&gt;
&lt;p&gt;&lt;em&gt;So now we unleash the beast. Firing all 500 requests at the same time.&lt;/em&gt;&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;finished in 16.66s, 30.02 req/s, 3.55MB/s&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;requests: 500 total, 500 started, 500 done, 500 succeeded, 0 failed, 0 errored, 0 timeout&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;CPU on API still doing good. A slight hint of CPU throttling due to CFS, which is used when you set CPU limits in Kubernetes.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./3-burst-http2-500-concurrent-analysis-cpu.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Important about Kubernetes and CPU limits&lt;br /&gt;
Even with CPU limits set to 1 (100% of one CPU), your container can still be throttled at much lower CPU usage. Check out &lt;a href=&quot;https://medium.com/omio-engineering/cpu-limits-and-aggressive-throttling-in-kubernetes-c5b20bd8a718&quot;&gt;this article&lt;/a&gt; for more information.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Whopsie. The average database query latency has increased drastically, and we have a long tail of very slow queries. Looks like we are starting to see signs of bottlenecks on the database. This might also be affected by our maximum of 60 concurrent connections to the database, resulting in queries having to wait their turn before executing.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./3-burst-http2-500-concurrent-analysis-db-latency.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;It’s hard to judge the peak rate of database queries due to limited sampling of the metrics.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./3-burst-http2-500-concurrent-analysis-db-queries.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Now individual HTTP requests are much slower due to waiting for the database.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./3-burst-http2-500-concurrent-analysis-http-latency.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Here is just a random trace from Elastic APM to see if the increased database latency is concentrated to specific queries or tables or just general saturation. Indeed there is a single query responsible for half the time taken for the entire query! We better get back to that in a bit and dig further.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./3-burst-http2-500-concurrent-analysis-apm-trace.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;In an ideal world all 500 requests should start and complete in 2-300ms regardless. Since that is not happening it’s an indication that we are now hitting some other bottleneck.&lt;/p&gt;
&lt;p&gt;Looking at the graphs it seems we are starting to saturate the database. The latency for every request is now largely dependent on the slowest of the 10-12 database queries it depends on. And as we are stressing the database the probability of slow queries increase. The latency for the whole process of fetching 500 requests are again largely dependent on the slowest requests.&lt;/p&gt;
&lt;p&gt;So this optimization gives on average better performance, but more variability of the individual requests, when the system is under heavy load.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;DatabaseOptimizations&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;side-quest-database-optimizations&quot;&gt;Side quest: Database optimizations&lt;/h2&gt;
&lt;p&gt;It seems we are saturating the database. Before throwing more money at the problem (by increasing database size) I like to know what the bottlenecks are. Looking at the traces from APM
I see one query that is consistently taking 10x longer than the rest. I also confirm this in the AWS RDS Performance Insights that show the top SQL queries by load.&lt;/p&gt;
&lt;p&gt;When designing the database schema I came up with the idea of having immutability for certain data types. So instead of overwriting row with ID 1, we add a row with ID 1 Revision 2. Now we have the history of who did what to the data and can easily track changes and roll back if needed. The most common use case is just fetching the last revision. So for simplicity I created a PostgreSQL view that only shows the last revision. That way clients don’t have to worry about the existense of revisions at all. That is now just an implementation detail.&lt;/p&gt;
&lt;p&gt;When it comes to performance that turns out to be an important implementation detail. The view is using &lt;code&gt;SELECT DISTINCT ON (id) ... ORDER BY id, revision DESC&lt;/code&gt;. However many of the queries to the view is ordering the returned data by time, and expect the data returned from database to already be ordered chronologically. Using &lt;code&gt;EXPLAIN ANALYZE&lt;/code&gt; on the queries this always results in a full table scan instead of using indexes, and is what’s causing this specific query to be slow. Without going into details it seems there is no simple and efficient way of having a view with the last revision and query that for a subset of rows ordered again by time.&lt;/p&gt;
&lt;p&gt;For the forseable future this does not actually impact real world usage. It’s only apparent under artificially large loads under the worst conditions. But now we know where we need to refactor things if performance actually becomes a problem.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;NextBottleneck&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;determining-the-next-bottleneck&quot;&gt;Determining the next bottleneck&lt;/h2&gt;
&lt;p&gt;Whenever I fix one problem I like to know where, how and when the next problem or limit is likely to appear. When increasing the number of requests and streams I expected to see increasing latency. But instead I see errors appear like a cliff:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;finished in 27.33s, 36.59 req/s, 5.64MB/s&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;requests: 5000 total, 1002 started, 1002 done, 998 succeeded, 4002 failed, 4000 errored, 0 timeout&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Consulting the logs for both the nginx load balancer and the API there are no records of failing requests. Since nginx does not pass the HTTP2 connection directly to the API, but instead “unbundles” them into HTTP1 requests I suspect there might be issues with connection limits or even available ports from nginx to the API. But maybe it’s a configuration issue. By default nginx does &lt;a href=&quot;http://nginx.org/en/docs/http/ngx_http_upstream_module.html#server&quot;&gt;not limit the number of connections to a backend&lt;/a&gt; (our API). . But, there is actually a &lt;a href=&quot;https://nginx.org/en/docs/http/ngx_http_v2_module.html#http2_max_requests&quot;&gt;default limit to the number of HTTP2 requests that can be served over a single connection&lt;/a&gt; - And it happens to be 1000.&lt;/p&gt;
&lt;p&gt;I leave it at that. It’s very unlikely we’ll be hitting these limits any time soon.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;ClusterResources&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;side-quest-cluster-resources-and-burstable-vms&quot;&gt;Side quest: Cluster resources and burstable VMs&lt;/h2&gt;
&lt;p&gt;When load testing the first time around sometimes Grafana would also become unresponsive. That’s usually a bad sign. It might indicate that the underlying infrastructure is also reaching saturation. That is not good since it can impact what should be independent services.&lt;/p&gt;
&lt;p&gt;Our Kubernetes cluster is composed of 2x t3a.medium on demand nodes and 2x t3a.medium spot nodes. These VM types are burstable. You can use 20% per vCPU sustained over time without problems. If you exceed those 20% you start consuming CPU credits faster than they are granted and once you run out of CPU credits processes will be forcibly throttled.&lt;/p&gt;
&lt;p&gt;Of course Kubernetes does not know about this and expects 1 CPU to actually be 1 CPU. In addition Kubernetes will decide where to place workloads based on their stated resource requirements and limits, and not their actual resource usage.&lt;/p&gt;
&lt;p&gt;When looking at the actual metrics two of our nodes are indeed out of CPU credits and being throttled. The sum of factors leading to this is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;We have not yet set resource requests and limits making it harder for Kubernetes to intelligently place workloads&lt;/li&gt;
&lt;li&gt;Using burstable nodes having some additional constraints not visible to Kubernetes&lt;/li&gt;
&lt;li&gt;Old deployments laying around consuming unnecessary resources&lt;/li&gt;
&lt;li&gt;Adding costly features without assessing the overall impact&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I have not touched on the last point yet. I started adding &lt;a href=&quot;https://pyroscope.io/&quot;&gt;Pyroscope&lt;/a&gt; to our systems since I simply love monitoring All The Things. The documentation does not go into specifics but emphasizes that it’s “low overhead”. Remember that our budget for CPU usage is actually 40% per node, not 200%. The Pyroscope server itself consumes 10-15% CPU which seems fair. But investigating further the Pyroscope agent also consumes 5-6% CPU per instance. This graph shows the CPU usage of a single Pod before and after turning off Pyroscope profiling.&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;pyroscope-agent-cpu&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./pyroscope-agent-cpu.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;5-6% CPU overhead on a highly utilized service is probably worth it. But when the baseline CPU usage is 0% CPU and we have multiple services and deployments in different environments we are suddenly using 40-60% CPU on profiling and less than 1% on actual work!&lt;/p&gt;
&lt;p&gt;The outcome of this is that we need to separate burstable and stable load deployments. Monitoring and supporting systems are usually more stable resource wise while the actual business systems much more variable, and suitable for burst nodes. In practice we add a node pool of non-burst VMs and use NodeAffinity to stick Prometheus, Pyroscope etc to those nodes. Another benefit of this is that the supporting systems needed to troubleshoot problems are now less likely to be impacted by the problem itself, making troubleshooting much easier.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Conclusion&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;This whole adventure only took a few hours but resulted in some specific and immediate performance gains. It also highlighted the weakest links in our application, database and infrastructure architecture.&lt;/p&gt;
</content:encoded><category>tech</category><category>kubernetes</category><category>databases</category><category>observability</category></item><item><title>End of 2020 rough database landscape</title><link>https://blog.stian.omg.lol/p/end-of-2020-rough-database-landscape/</link><guid isPermaLink="true">https://blog.stian.omg.lol/p/end-of-2020-rough-database-landscape/</guid><description>This post is an attempt to lay out the rough landscape of databases that you might encounter or consider as of late 2020.</description><pubDate>Fri, 27 Nov 2020 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://blog.stian.omg.lol/_astro/2020-11-27-end-of-2020-rough-database-landscape.CNYs8WYf_Z1is9QE.jpg&quot; width=&quot;1600&quot; height=&quot;914&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;This post is an attempt to lay out the rough landscape of databases that you might encounter or consider as of late 2020.&lt;/p&gt;
&lt;p&gt;There seems to exist a database for every niche, mood or emotion. And they seem to change just as fast.&lt;/p&gt;
&lt;p&gt;How do you balance the urge for the new and shiny but without risking too much headache down the road?&lt;/p&gt;
&lt;p&gt;This post is an attempt to lay out the rough landscape of databases that you might encounter or consider as of late 2020.&lt;/p&gt;
&lt;p&gt;There will be broad generalizations for brevity.&lt;/p&gt;
&lt;p&gt;The goal is not to be exhaustive or take all possible precautions. Consider it a starting point for further research and planning.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;TLDR: Scroll to the &lt;a href=&quot;#Landscape&quot;&gt;diagrams&lt;/a&gt; or view the &lt;a href=&quot;https://blog.stian.omg.lol/attachments/2020-11-27-end-of-2020-rough-database-landscape/map-complete.png&quot;&gt;big picture&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Table of contents&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#Background&quot;&gt;Background&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#ProjectPhase&quot;&gt;Project phase overview&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Planning&quot;&gt;Planning&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#DatabaseCategories&quot;&gt;Database categories&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#SQL&quot;&gt;SQL&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#NoSQL&quot;&gt;NoSQL&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#KeyValue&quot;&gt;KeyValue&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Timeseries&quot;&gt;Timeseries&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Graph&quot;&gt;Graph&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#OtherNiceThings&quot;&gt;Other nice things&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Landscape&quot;&gt;The Landscape&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#SQLMap&quot;&gt;SQL&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#NoSQLMap&quot;&gt;NoSQL&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#KeyValueMap&quot;&gt;KeyValue&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#TimeseriesMap&quot;&gt;Timeseries&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#GraphMap&quot;&gt;Graph&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#FurtherReading&quot;&gt;Further reading&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Conclusion&quot;&gt;Conclusion&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a id=&quot;Background&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;background&quot;&gt;Background&lt;/h2&gt;
&lt;p&gt;I’m a consultant doing development, DevOps and cloud infrastructure. I also have the occasional side project trying out the Tech Flavor of the Month.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;ProjectPhase&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;project-phase-overview&quot;&gt;Project phase overview&lt;/h3&gt;
&lt;p&gt;The typical phases in projects I’m involved in follow no scientific or trademarked methodology, so YMMV:&lt;/p&gt;
&lt;h4 id=&quot;starting-out&quot;&gt;Starting out&lt;/h4&gt;
&lt;p&gt;Get something working as fast as possible. Take all the shortcuts. Use some opinionated framework or platform.&lt;/p&gt;
&lt;h4 id=&quot;moving-from-development-to-production&quot;&gt;Moving from development to production&lt;/h4&gt;
&lt;p&gt;People like it, people use it. Move the thing from a single “pet server” to a more robust cloud environment.&lt;/p&gt;
&lt;h4 id=&quot;scaling-production&quot;&gt;Scaling production&lt;/h4&gt;
&lt;p&gt;Bottlenecks and scaling problems start to emerge. Refactor or replace some pieces to remove the bottlenecks.&lt;/p&gt;
&lt;h4 id=&quot;challenges&quot;&gt;Challenges&lt;/h4&gt;
&lt;p&gt;Moving between these phases might be a major PITA if the wrong shortcuts were taken in the previous phases.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;This of course applies to all technology choices and not just databases. But we have to start somewhere, right?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Planning&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;planning&quot;&gt;Planning&lt;/h2&gt;
&lt;p&gt;When starting out I try to envision all the phases of the project and which directions it may take in the future.&lt;/p&gt;
&lt;p&gt;First I want the technology or software I choose to be instantly usable. A Docker image. Great. An &lt;code&gt;apt-get install&lt;/code&gt;. Sweet. &lt;code&gt;npm install&lt;/code&gt;. Sure, why not. Downloading a tarball. Installing some C dependencies. Setting some flags. Compiling. Symlinking and fixing permissions. Creating some configuration from scratch. Making my own systemd service definitions. Going back and doing every step again because it failed. &lt;em&gt;Mkay, no thanks, I’m out.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;At least for me it’s a plus if it’s easy to deploy on Kubernetes since I use it for everything already. I always have a cluster or three laying around so I can get a prototype or five up and running quickly before later spending money for cloud hosting.&lt;/p&gt;
&lt;p&gt;Does the thing have momentum and a community? If it does it probably has high quality tooling either by the vendor or the open source community (preferably both). It probably also has lots of common questions answered on blogs and StackOverflow and Github issues.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;So we managed to build something and the audience likes it.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;How easy is it to move it from a production environment into something stable and low-maintenance? For databases that would typically involve using a managed service for hosting it. You do not want to be responsible for operating your own databases. Is it common enough that there are competitors in the marketplace offering it as a managed service? If there is only a single option expect prices to be very steep. Preferably also a managed service by one of the big known cloud platforms. They are usually cheaper. They are less likely to vanish. It might make integration with other systems easier later.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;We hit some problems either because of raw scale or some type of usage we did not anticipate in the beginning.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Are there compatible implementations that might solve some common problems? Typically this is because an implementation has to make a decision about it’s trade-offs. For a database system this is usually around the CAP theorem. A database system (or anything that keeps state) can be:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;Partition Tolerant&lt;/em&gt; - The system still works if a node or the network between nodes fail.&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Available&lt;/em&gt; - All requests receive a response.&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Consistent&lt;/em&gt; - The data we read is the current data and not an earlier state.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But, you can only have two at the same time. And distributed systems tends to need to be partition tolerant. So we are stuck between consistency and availability.&lt;/p&gt;
&lt;p&gt;It might be a good to have an idea of the CAP tradeoffs an implementation has done, and whether there are compatible implementations with different tradeoffs that can be used if later we find out we need to tweak our trade-offs for speed and/or scale.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;More information about CAP theorem &lt;a href=&quot;https://en.wikipedia.org/wiki/CAP_theorem&quot;&gt;here&lt;/a&gt; and &lt;a href=&quot;https://towardsdatascience.com/cap-theorem-and-distributed-database-management-systems-5c2be977950e&quot;&gt;here&lt;/a&gt;. Jepsen have also &lt;a href=&quot;https://jepsen.io/analyses&quot;&gt;extensively tested&lt;/a&gt; many popular databases to see how they break and if they are true to their stated trade-offs.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;DatabaseCategories&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;database-categories&quot;&gt;Database categories&lt;/h3&gt;
&lt;p&gt;Databases can be roughly sorted into categories. I’ll keep it simple and use the everyday lingo and not go into details about semantics and definitions (forgive me).&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://www.prisma.io/dataguide/intro/comparing-database-types&quot;&gt;https://www.prisma.io/dataguide/intro/comparing-database-types&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;SQL&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;sql&quot;&gt;SQL&lt;/h4&gt;
&lt;p&gt;The oldest category is the relational database, also known as SQL based on the typical interface used to access these databases.&lt;/p&gt;
&lt;p&gt;In general these databases have tables with names, a set of pre-defined columns and an arbitrary number of rows. You should have an idea of the data types to be stored in each column (such as text or numbers).&lt;/p&gt;
&lt;p&gt;The downside of this is that you have to start with a rough model of the data you want to store and work with. The benefit of this is that later you know something about the model of the data you are working with. Most of the time I’ll happily do this in the database rather than handle all the potential inconsistencies in all systems that use that database.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Main contenders: PostgreSQL. MySQL &amp;amp; MariaDB.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;NoSQL&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;nosql&quot;&gt;NoSQL&lt;/h4&gt;
&lt;p&gt;All the rage the last decade. You put data in you get data out. The data is structured but not necessarily predefined. Think JSON object with values, arrays and lists.&lt;/p&gt;
&lt;p&gt;The benefit is productivity when developing. The drawback is that you may pay a price for those shortcuts later if you’re not careful.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Main contender: MongoDB.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;KeyValue&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;keyvalue&quot;&gt;KeyValue&lt;/h4&gt;
&lt;p&gt;Technically a sub-category of NoSQL, and should probably be called caches. But I feel it deserves it’s own category.&lt;/p&gt;
&lt;p&gt;A hyper-fast hyper-simple type of database. It has two columns. A key (ID) and value. The value can be anything, a string, a number, an entire JSON object or a blob containing binary data.&lt;/p&gt;
&lt;p&gt;These are typically used in combination with another type of database. Either by storing very commonly used data for even quicker access. Or for certain types of simple data that requires insane speed or throughput and you don’t want to overload the main database.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Main contender: Redis.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Timeseries&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;timeseries&quot;&gt;Timeseries&lt;/h4&gt;
&lt;p&gt;A lesser known type of database optimized for storing a time series. A time series is a specific data type where the index is typically the time of a measurement. And the measurement is a number.&lt;/p&gt;
&lt;p&gt;A time series is almost never changed after the fact. So these databases can be optimized for writing huge amounts of new data and reading and calculating on existing data. At the cost of performance for deleting or updating old data which is sloooow. Since the values are always numbers that tend to change somewhat predictably compression and deduplication can save us massive amounts of storage.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Main contenders: Prometheus, InfluxDB, TimescaleDB (plugin for PostgreSQL).&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Graph&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;graph&quot;&gt;Graph&lt;/h4&gt;
&lt;p&gt;Graph databases are cool. In a graph database the relationship between objects is a primary feature. Whereas in SQL you need to join an element from one table with another object in another table with some kind of common identifier.&lt;/p&gt;
&lt;p&gt;For most simple use cases a regular SQL database will do fine. But when the number of objects stored (rows) and the number of intermediary tables (joins) become large it gets slow, or expensive, or both.&lt;/p&gt;
&lt;p&gt;I don’t have much experience with graph databases but I suspect they are less suited to general tasks and should be reserved for solving specific problems.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Main contenders: Neo4j. Redis + RedisGraph.&lt;/em&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;PS: Graph databases and GraphQL are completely separate things.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;OtherNiceThings&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;other-nice-things&quot;&gt;Other nice things&lt;/h4&gt;
&lt;p&gt;When researching this post I’ve come across things that look promising but are hard to categorize or fall in their own very niche categories.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://dgraph.io&quot;&gt;Dgraph&lt;/a&gt; - A GraphQL and backend in one.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://prestodb.io&quot;&gt;PrestoDB&lt;/a&gt; - An SQL interface on top of whatever database or storage you want to connect.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://rethinkdb.com&quot;&gt;RethinkDB&lt;/a&gt; - A NoSQL database focused on real-time streaming/updating clients.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.foundationdb.org&quot;&gt;FoundationDB&lt;/a&gt; - A transactional key-value store by Apple.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clickhouse.tech/&quot;&gt;ClickHouse&lt;/a&gt; - An SQL database that stores data (on disk) in columns instead of rows. Makes for blazingly fast analytical and aggregation queries.&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://aws.amazon.com/qldb/&quot;&gt;Amazon Quantum Ledger Database&lt;/a&gt; - A managed distributed ledger database (aka blockchain).&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.enterprisedb.com/products/edb-postgres-advanced-server-secure-ha-oracle-compatible&quot;&gt;EDB Postgres Advanced Server&lt;/a&gt; - An Oracle compatible PostgreSQL variant.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a id=&quot;Landscape&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;the-landscape&quot;&gt;The Landscape&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;How to use these maps:&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Version compatibility are in parenthesis. I have not mapped every version and how much breaking they are compared to previous versions but included some notes where I know there might be issues.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;API/Protocol/Interface&lt;/em&gt; - This is decided by the framework, tool or driver you want to use. Sometimes it might be easier to choose the framework first and then a fitting database protocol. Or you might be lucky to choose the database features you need first and then select frameworks, tools and drivers that support it.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I think interfaces are really important when creating and choosing technology. I had a &lt;a href=&quot;https://speakerdeck.com/stianovrevage/avoiding-lock-in-without-avoiding-managed-services&quot;&gt;presentation&lt;/a&gt; about it a while ago and I think it’s still relevant.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;em&gt;Engine&lt;/em&gt; - Database implementations that are independent but try to be compatible. If there are alternatives to the “original” implementation they might have done different tradeoffs with regards to the CAP theorem or solve other specific problems.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Big three managed&lt;/em&gt; - Available managed services by the big three clouds, Amazon (AWS), Google (GCP) or Microsoft (Azure). Having an option to host in the big three is most likely the cheapest method as well as having a variety of other managed services to build a complete system in a single cloud.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Vendor managed&lt;/em&gt; - If the database vendor or backing company offers an Official managed service. They are usually hosted on the big three. Potentially a large cost premium over the raw compute power.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Self-hosted&lt;/em&gt; - Implementations you can run on your own computer or server.&lt;/p&gt;
&lt;p&gt;| Legend  |
|––––|—–|
| &lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;icon-checklist&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./icon-checklist.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt; | The checklist icon marks potential compatibility issues. For most use cases not a problem.&lt;br&gt;&lt;strong&gt;PS:&lt;/strong&gt; The absence of this icon does not automatically mean compatibility. |
| &lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;icon-operator&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./icon-operator.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt; | I put the lightning icon on the self-hosted implementations that have what seems to be stable Kubernetes operators available. In short, a Kubernetes operator makes running a stateful system, such as a database, on Kubernetes much easier. It might allow for longer time before migrating to a managed service. |&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;SQLMap&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;sql-1&quot;&gt;SQL&lt;/h3&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;map-sql&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./map-sql.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Compatibility:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://blog.yugabyte.com/postgresql-compatibility-in-yugabyte-db-2-0/&quot;&gt;PostgreSQL - Yugabyte&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.cockroachlabs.com/docs/stable/postgresql-compatibility.html&quot;&gt;PostgreSQL - CockroachDB&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://mariadb.com/kb/en/mariadb-vs-mysql-compatibility/&quot;&gt;MySQL - MariaDB&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;Kubernetes Operators:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/CrunchyData/postgres-operator&quot;&gt;PostgreSQL (CrunchyData)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/zalando/postgres-operator&quot;&gt;PostgreSQL (Zalando)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.yugabyte.com/latest/deploy/kubernetes/single-zone/oss/yugabyte-operator/&quot;&gt;Yugabyte&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/cockroachdb/cockroach-operator&quot;&gt;CockroachDB&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.percona.com/software/percona-kubernetes-operators&quot;&gt;Percona PostgreSQL for MySQL &amp;amp; XtraDB&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;NoSQLMap&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;nosql-1&quot;&gt;NoSQL&lt;/h3&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;map-nosql&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./map-nosql.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;PS: There are some &lt;a href=&quot;https://docs.mongodb.com/manual/release-notes/4.0-compatibility/&quot;&gt;breaking changes&lt;/a&gt; from MongoDB 3.6 to 4 so make sure the tools you intend to use are compatible with the database version you intend on using.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;Kubernetes Operators:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/mongodb/mongodb-kubernetes-operator&quot;&gt;MongoDB&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.percona.com/doc/kubernetes-operator-for-psmongodb/index.html&quot;&gt;Percona Distribution for MongoDB&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/scylladb/scylla-operator&quot;&gt;ScyllaDB&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.elastic.co/guide/en/cloud-on-k8s/current/k8s-overview.html&quot;&gt;Elastic Stack&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;KeyValueMap&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;keyvalue-1&quot;&gt;KeyValue&lt;/h3&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;map-keyvalue&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./map-keyvalue.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Kubernetes Operators:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/spotahome/redis-operator&quot;&gt;Redis (Spotahome)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;TimeseriesMap&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;timeseries-1&quot;&gt;Timeseries&lt;/h3&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;map-timeseries&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./map-timeseries.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Kubernetes Operators:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/prometheus-community/helm-charts/tree/main/charts/kube-prometheus-stack&quot;&gt;Prometheus-Stack&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/VictoriaMetrics/operator&quot;&gt;VictoriaMetrics&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;GraphMap&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;graph-1&quot;&gt;Graph&lt;/h3&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;map-graph&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./map-graph.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Kubernetes Operators:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.arangodb.com/docs/stable/deployment-kubernetes-usage.html&quot;&gt;ArangoDB&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;FurtherReading&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;further-reading&quot;&gt;Further reading&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://en.wikipedia.org/wiki/Comparison_of_relational_database_management_systems&quot;&gt;Wikipedia on RDBMS&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://db-engines.com/en/&quot;&gt;DB-engines.com&lt;/a&gt; - Lots of statistics and comparisons between DB engines&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://landscape.cncf.io/&quot;&gt;CNCF Landscape&lt;/a&gt; - What’s moving in the cloud native landscape, including databases.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a id=&quot;Conclusion&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Congratulations if you made it this far!&lt;/p&gt;
&lt;p&gt;I did this research primarily to reduce my own analysis paralysis on various projects so I can get-back-to-building. If you learned something as well, great stuff!&lt;/p&gt;
&lt;p&gt;And if you want my advice, just use PostgreSQL unless you really know about some special requirements that necessitates using something else :-)&lt;/p&gt;
</content:encoded><category>tech</category><category>databases</category></item><item><title>Mini-post: Down-scaling Azure Kubernetes Service (AKS)</title><link>https://blog.stian.omg.lol/p/mini-post-down-scaling-azure-kubernetes-service-aks/</link><guid isPermaLink="true">https://blog.stian.omg.lol/p/mini-post-down-scaling-azure-kubernetes-service-aks/</guid><description>Looking into caveats of running Kubernetes on AKS with small node sizes and node counts.</description><pubDate>Tue, 04 Jun 2019 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://blog.stian.omg.lol/_astro/2019-06-04-downscaling-aks.DEagaLD__Z2k3IVc.jpg&quot; width=&quot;1600&quot; height=&quot;914&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;Looking into caveats of running Kubernetes on AKS with small node sizes and node counts.&lt;/p&gt;
&lt;p&gt;We discovered today that some implicit assumptions we had about AKS at smaller scales were incorrect.&lt;/p&gt;
&lt;p&gt;Suddenly new workloads and jobs in our Radix CI/CD could not start due to insufficient resources (CPU &amp;amp; memory).&lt;/p&gt;
&lt;p&gt;Even though it only caused problems in development environments with smaller node sizes it still surprised some of our developers, since we expected the size of development clusters to have enough resources.&lt;/p&gt;
&lt;p&gt;I thought it would be a good chance to go a bit deeper and verify some of our assumptions and also learn more about various components that usually “just works” and isn’t really given much thought.
First I do a &lt;code&gt;kubectl describe node &amp;lt;node&amp;gt;&lt;/code&gt; on 2-3 of the nodes to get an idea of how things are looking:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;sh&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;Resource&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;                       Requests&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;          Limits&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;--------&lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;                       --------&lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;          ------&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;cpu&lt;/span&gt;&lt;span style=&quot;--shiki-light:#032F62;--shiki-dark:#9ECBFF&quot;&gt;                            930m&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; (98%)        5500m (&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;585%&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;memory&lt;/span&gt;&lt;span style=&quot;--shiki-light:#005CC5;--shiki-dark:#79B8FF&quot;&gt;                         1659939584&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt; (89%)  4250M (&lt;/span&gt;&lt;span style=&quot;--shiki-light:#6F42C1;--shiki-dark:#B392F0&quot;&gt;228%&lt;/span&gt;&lt;span style=&quot;--shiki-light:#24292E;--shiki-dark:#E1E4E8&quot;&gt;)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;So we are obviously hitting the roof when it comes to resources. But why?&lt;/p&gt;
&lt;h3 id=&quot;node-overhead&quot;&gt;Node overhead&lt;/h3&gt;
&lt;p&gt;We use &lt;code&gt;Standard DS1 v2&lt;/code&gt; instances as AKS nodes and they have 1 CPU core and 3.5 GiB memory.&lt;/p&gt;
&lt;p&gt;The output of &lt;code&gt;kubectl describe node&lt;/code&gt; also gives us info on the Capacity (total node size) and Allocatable (resources available to run Pods).&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Capacity:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt; cpu:                            1&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt; memory:                         3500452Ki&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Allocatable:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt; cpu:                            940m&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt; memory:                         1814948Ki&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;So we have lost &lt;strong&gt;60 millicores / 6%&lt;/strong&gt; of CPU and &lt;strong&gt;1685MiB / 48%&lt;/strong&gt; of memory. The next question is if this increases linearly with node size (the percentage of resources lost is the same regardless of node size) or is fixed (always reserves 60 millicores and 1685Mi of memory), or a combination.&lt;/p&gt;
&lt;p&gt;I connect to another cluster that has double the node size (&lt;code&gt;Standard DS2 v2&lt;/code&gt;) and compare:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Capacity:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt; cpu:                            2&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt; memory:                         7113160Ki&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Allocatable:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt; cpu:                            1931m&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt; memory:                         4667848Ki&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;So for this the loss is &lt;strong&gt;69 millicores / 3.5%&lt;/strong&gt; of CPU and &lt;strong&gt;2445MiB / 35%&lt;/strong&gt; of memory.&lt;/p&gt;
&lt;p&gt;So CPU reservations are close to fixed regardless of node size while memory reservations are influenced by node size but luckily not linearly.&lt;/p&gt;
&lt;p&gt;What causes this “waste”? Reading up on &lt;a href=&quot;https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/&quot;&gt;kubernetes.io&lt;/a&gt; gives a few clues. Kubelet will reserve CPU and memory resources for itself and other Kubernetes processes. It will also reserve a portion of memory to act as a buffer whenever a Pod is going beyond it’s memory limits to avoid risking System OOM, potentially making the whole node unstable.&lt;/p&gt;
&lt;p&gt;To figure out what these are configured to we log in to an actual AKS node’s console and run &lt;code&gt;ps ax|grep kube&lt;/code&gt; and the output looks like this:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;/usr/local/bin/kubelet --enable-server --node-labels=node-role.kubernetes.io/agent=,kubernetes.io/role=agent,agentpool=nodepool1,storageprofile=managed,storagetier=Premium_LRS,kubernetes.azure.com/cluster=MC_clusters_weekly-22_northeurope --v=2 --volume-plugin-dir=/etc/kubernetes/volumeplugins --address=0.0.0.0 --allow-privileged=true --anonymous-auth=false --authorization-mode=Webhook --azure-container-registry-config=/etc/kubernetes/azure.json --cgroups-per-qos=true --client-ca-file=/etc/kubernetes/certs/ca.crt --cloud-config=/etc/kubernetes/azure.json --cloud-provider=azure --cluster-dns=10.2.0.10 --cluster-domain=cluster.local --enforce-node-allocatable=pods --event-qps=0 --eviction-hard=memory.available&amp;lt;750Mi,nodefs.available&amp;lt;10%,nodefs.inodesFree&amp;lt;5% --feature-gates=PodPriority=true,RotateKubeletServerCertificate=true --image-gc-high-threshold=85 --image-gc-low-threshold=80 --image-pull-progress-deadline=30m --keep-terminated-pod-volumes=false --kube-reserved=cpu=60m,memory=896Mi --kubeconfig=/var/lib/kubelet/kubeconfig --max-pods=110 --network-plugin=cni --node-status-update-frequency=10s --non-masquerade-cidr=0.0.0.0/0 --pod-infra-container-image=k8s.gcr.io/pause-amd64:3.1 --pod-manifest-path=/etc/kubernetes/manifests --pod-max-pids=-1 --rotate-certificates=false --streaming-connection-idle-timeout=5m&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;To log in to the console of a node, go to the MC_resourcegroup_clustername_region resource-group and select the VM. Then go to &lt;code&gt;Boot diagnostics&lt;/code&gt; and enable it. Go to &lt;code&gt;Reset password&lt;/code&gt; to create yourself a user and then &lt;code&gt;Serial console&lt;/code&gt; to log in and execute commands.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;We can see &lt;code&gt;--kube-reserved=cpu=60m,memory=896Mi&lt;/code&gt; and &lt;code&gt;--eviction-hard=memory.available&amp;lt;750Mi&lt;/code&gt; which adds up to &lt;code&gt;1646Mi&lt;/code&gt; which is pretty close to the &lt;code&gt;1685Mi&lt;/code&gt; that was the gap between Capacity and Allocatable.&lt;/p&gt;
&lt;p&gt;We also do this on a &lt;code&gt;Standard DS2 v2&lt;/code&gt; node and get &lt;code&gt;--kube-reserved=cpu=69m,memory=1638Mi&lt;/code&gt; and &lt;code&gt;--eviction-hard=memory.available&amp;lt;750Mi&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;So we can see that the memory of &lt;code&gt;kube-reserved&lt;/code&gt; grows almost linearly and seems to always be about 20-25% while CPU reservations are almost the same. The memory eviction buffer is always fixed at &lt;code&gt;750Mi&lt;/code&gt; which would mean bigger resource waste as nodes decrease in size.&lt;/p&gt;
&lt;h5 id=&quot;cpu&quot;&gt;CPU&lt;/h5&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;Standard DS1 v2&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;Standard DS2 v2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;VM capacity&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;1.000m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;2.000m&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;kube-reserved&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;-60m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;-69m&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Allocatable&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;940m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;1.931m&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Allocatable %&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;94%&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;96.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h5 id=&quot;memory&quot;&gt;Memory&lt;/h5&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;Standard DS1 v2&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;Standard DS2 v2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;VM capacity&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;3.500Mi&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;7.113Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;kube-reserved&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;-896Mi&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;-1.638Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Eviction buf&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;-750Mi&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;-750Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Allocatable&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;1.814Mi&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;4.667Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Allocatable %&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;52%&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;65%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id=&quot;node-pods-daemonsets&quot;&gt;Node pods (DaemonSets)&lt;/h3&gt;
&lt;p&gt;We have some Pods that run on every node, and they are installed by default by AKS. We get the resource limits of these by describing either the pods or the daemonsets.&lt;/p&gt;
&lt;h5 id=&quot;cpu-1&quot;&gt;CPU&lt;/h5&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;Standard DS1 v2&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;Standard DS2 v2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Allocatable&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;940m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;1.931m&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;kube-system/calico-node&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;-250m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;-250m&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;kube-system/kube-proxy&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;-100m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;-100m&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;kube-system/kube-svc-redirect&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;-5m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;-5m&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Available&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;585m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;1.576m&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Available %&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;58%&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;81%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h5 id=&quot;memory-1&quot;&gt;Memory&lt;/h5&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;Standard DS1 v2&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;Standard DS2 v2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Allocatable&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;1.814Mi&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;4.667Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;kube-system/kube-svc-redirect&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;-32Mi&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;-32Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Available&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;1.782Mi&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;4.635Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Available %&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;50%&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;61%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;So for &lt;code&gt;Standard DS1 v2&lt;/code&gt; nodes we have about 0.5 CPU and 1.7GiB memory per node for pods. And for &lt;code&gt;Standard DS2 v2&lt;/code&gt; nodes it’s about 1.5 CPU and 4.6GiB memory.&lt;/p&gt;
&lt;h3 id=&quot;kube-system-pods&quot;&gt;kube-system pods&lt;/h3&gt;
&lt;p&gt;Now lets add some standard Kubernetes pods we need to run. As far as I know these are pretty much fixed for a cluster and not related to node size or count.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Deployment&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;CPU&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;Memory&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;kube-system/kubernetes-dashboard&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;100m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;50Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;kube-system/tunnelfront&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;10m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;64Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;kube-system/coredns (x2)&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;200m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;140Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;kube-system/coredns-autoscaler&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;20m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;10Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;kube-system/heapster&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;130m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;230Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sum&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;460m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;494Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id=&quot;third-party-pods&quot;&gt;Third party pods&lt;/h3&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Deployment&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;CPU&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;Memory&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;grafana&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;200m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;500Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;prometheus-operator&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;500m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;1.000Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;prometheus-alertmanager&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;100m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;225Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;flux&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;50m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;64Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;flux-helm-operator&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;50m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;64Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sum&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;900m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;1.853Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id=&quot;radix-platform-pods&quot;&gt;Radix platform pods&lt;/h3&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Deployment&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;CPU&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;Memory&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;radix-api-prod/server (x2)&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;200m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;400Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;radix-api-qa/server (x2)&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;100m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;200Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;radix-canary-golang-dev/www&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;40m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;500Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;radix-canary-golang-prod/www&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;40m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;500Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;radix-platform-prod/public-site&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;5m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;10Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;radix-web-console-prod/web&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;10m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;42Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;radix-web-console-qa/web&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;5m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;21Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;radix-github-webhook-prod/webhook&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;10m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;30Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;radix-github-webhook-prod/webhook&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;5m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;15Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sum&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;415m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;1.718Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;If we add up the resource usage of these groups of workloads and see the total available resources on our 4 node Standard DS1 v2 clusters we are left with 0.56 CPU cores (14%) and 3GB of memory (22%):&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;CPU&lt;/th&gt;
&lt;th style=&quot;text-align: right&quot;&gt;Memory&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;kube-system&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;460m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;494Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;third-party&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;900m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;1.853Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;radix-platform&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;415m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;1.718Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sum&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;1.760m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;4.020Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Available on 4x DS1&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;2.340m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;7.128Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Available for workloads&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;565m&lt;/td&gt;
&lt;td style=&quot;text-align: right&quot;&gt;3.063Mi&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Though surprising that we lost this much resources before being able to deploy our actual customer applications, it should still be a bit of headroom.&lt;/p&gt;
&lt;p&gt;Going further I checked the resource requests on 8 customer pods deployed in 4 environments (namespaces). Even though none of them had a resource configuration in their &lt;code&gt;radixconfig.yaml&lt;/code&gt; files they still had resource requests and limits. Not surprising since we use LimitRange to set default resource requests and limits. The surprise was that half of them had 50Mi of memory and the other half 500Mi, seemingly at random.&lt;/p&gt;
&lt;p&gt;It turns out that we did an update to the LimitRange values a few days ago but that only applies to new Pods, so depending on if the Pods got re-created for any reason they may or may not have the old request of 500Mi, which in our case of small clusters will quickly drain the available resources.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Read more about LimitRange here: &lt;a href=&quot;https://kubernetes.io/docs/tasks/administer-cluster/manage-resources/memory-default-namespace/&quot;&gt;kubernetes.io&lt;/a&gt; , and here is the commit that eventually trickled down to reduce memory usage: &lt;a href=&quot;https://github.com/equinor/radix-operator/commit/f022fcde993efdf6cbcafb2c6632707a823a2a27&quot;&gt;github.com&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3 id=&quot;pod-scheduling&quot;&gt;Pod scheduling&lt;/h3&gt;
&lt;p&gt;Depending on the weight between CPU and memory requests and how often things get destroyed and re-created you may find yourself in a situation where you have enough resources in your cluster but new workloads are still Pending. This can happen when one resource type (e.g. CPU) is filled before another (e.g. memory), leading one or more resources to be stranded and unlikely to be utilized.&lt;/p&gt;
&lt;p&gt;Imagine for example a cluster that is already utilized like this:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;CPU&lt;/th&gt;
&lt;th&gt;Memory&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;node0&lt;/td&gt;
&lt;td&gt;94%&lt;/td&gt;
&lt;td&gt;86%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;node1&lt;/td&gt;
&lt;td&gt;80%&lt;/td&gt;
&lt;td&gt;89%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;node2&lt;/td&gt;
&lt;td&gt;98%&lt;/td&gt;
&lt;td&gt;60%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Scheduling a workload that requests 15% CPU and 20% memory cannot be scheduled since there are no nodes fulfilling both requirements. In theory there is probably a CPU intensive Pod on node2 that could be moved to node1 but Kubernetes does not do re-scheduling to optimize utilization. It can do re-scheduling based on Pod priority (&lt;a href=&quot;https://medium.com/@dominik.tornow/the-kubernetes-scheduler-cd429abac02f&quot;&gt;medium.com&lt;/a&gt;) and there is an incubator project (&lt;a href=&quot;https://akomljen.com/meet-a-kubernetes-descheduler/&quot;&gt;akomljen.com&lt;/a&gt;) that can try to drain nodes with low utilization.&lt;/p&gt;
&lt;p&gt;So for the foreseable future keeping in mind that resources can get stranded and that looking at the sum of cluster resources and sum of cluster resource demand might be misleading.&lt;/p&gt;
&lt;h3 id=&quot;calico-node&quot;&gt;calico-node&lt;/h3&gt;
&lt;p&gt;The biggest source of waste on our small clusters is &lt;code&gt;calico-node&lt;/code&gt; which is installed on every node and requests 25% of a CPU core while only using 2.5-3% CPU:&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;calico-node cpu usage&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./calico-node-cpu.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;The request is originally set here &lt;a href=&quot;https://github.com/Azure/aks-engine/blob/master/parts/k8s/containeraddons/kubernetesmasteraddons-calico-daemonset.yaml&quot;&gt;github.com&lt;/a&gt; but I have not got into why that number was choosen. Next steps would be to do some benchmarking of &lt;code&gt;calico-node&lt;/code&gt; to smoke out it’s performance characteristics to see if it would be safe to lower the resource requests, but that is out of scope for now.&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;By increasing node size from &lt;code&gt;Standard DS1 v2&lt;/code&gt; to &lt;code&gt;Standard DS2 v2&lt;/code&gt; we also increase the available CPU from 58% per node to 81% per node. Available memory increases from 50% to 61% per node.&lt;/li&gt;
&lt;li&gt;With a total platform requirement of 3-4GB of memory and 4.6GB available on &lt;code&gt;Standard DS2 v2&lt;/code&gt; we might have more resources for actual workloads on a 1-node &lt;code&gt;Standard DS2 v2&lt;/code&gt; cluster than a 3-node &lt;code&gt;Standard DS1 v2&lt;/code&gt; cluster!&lt;/li&gt;
&lt;li&gt;Beware of stranded resources limiting the utilization you can achieve across a cluster.&lt;/li&gt;
&lt;/ul&gt;
</content:encoded><category>tech</category><category>kubernetes</category><category>azure</category></item><item><title>Disk performance on Azure Kubernetes Service (AKS) - Part 1: Benchmarking</title><link>https://blog.stian.omg.lol/p/disk-performance-on-azure-kubernetes-service-aks-part-1-benchmarking/</link><guid isPermaLink="true">https://blog.stian.omg.lol/p/disk-performance-on-azure-kubernetes-service-aks-part-1-benchmarking/</guid><description>In this first post on troubleshooting some disk performance issues on Azure Kubernetes Service (AKS) we will benchmark Azure Premium SSD to find how workloads affect performance and which metrics to monitor to know when troubleshooting potential disk issues.</description><pubDate>Sat, 23 Feb 2019 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://blog.stian.omg.lol/_astro/2019-02-23-disk-performance-on-aks-part-1.D9zG2ory_Z2pHb0O.jpg&quot; width=&quot;1600&quot; height=&quot;914&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;In this first post on troubleshooting some disk performance issues on Azure Kubernetes Service (AKS) we will benchmark Azure Premium SSD to find how workloads affect performance and which metrics to monitor to know when troubleshooting potential disk issues.&lt;/p&gt;
&lt;p&gt;Understanding the characteristics of disk performance of a platform might be more important than you think. If disk resources are not correctly matched to your workload, your performance will suffer and might lead you to incorrectly diagnose a problem as being related to CPU or memory.&lt;/p&gt;
&lt;p&gt;The defaults might also not give you the performance you expect.&lt;/p&gt;
&lt;p&gt;In this first post on troubleshooting some disk performance issues on Azure Kubernetes Service (AKS) we will benchmark Azure Premium SSD to find how workloads affect performance and which metrics to monitor to know when troubleshooting potential disk issues.
TLDR:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Disable Azure cache for workloads with high number of random writes&lt;/li&gt;
&lt;li&gt;Use a P15 (256GB) or larger Premium SSD even though you might only need a fraction of it.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Table of contents&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#Background&quot;&gt;Background&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#MetricsMethodologies&quot;&gt;Metric Methodologies&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#StorageBackground&quot;&gt;Storage Background&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#WhatToMeasure&quot;&gt;What to measure?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#HowToMeasureDisk&quot;&gt;How to measure disk&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;HowToMeasureDiskOnAKS&quot;&gt;How to measure disk on Azure Kubernetes Service&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Tests&quot;&gt;Test results&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#Test1&quot;&gt;Test 1 - Learning to dislike Azure Cache&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Test2&quot;&gt;Test 2 - Disable Azure Cache - enable OS cache&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Test3&quot;&gt;Test 3 - Disable OS cache&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Test4&quot;&gt;Test 4 - Increase IO depth&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Test5&quot;&gt;Test 5 - Larger block size, smaller IO depth&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Test6&quot;&gt;Test 6 - Enable OS cache&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Test7&quot;&gt;Test 7 - Random writes, small block size&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Test8&quot;&gt;Test 8 - Large block size&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Conclusion&quot;&gt;Conclusion&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id=&quot;microsoft-azure&quot;&gt;Microsoft Azure&lt;/h4&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;a href=&quot;https://azure.microsoft.com/en-us/free/&quot;&gt;If you don’t have a Azure subscription already you can try services for $200 for 30 days.&lt;/a&gt; The VM size &lt;strong&gt;Standard_B2s&lt;/strong&gt; is Burstable, has 2vCPU, 4GB RAM, 8GB temp storage and costs roughly $38 / month. For $200 you can have a cluster of 3-4 B2s nodes plus traffic, loadbalancers and other additional costs.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;See my blog post &lt;a href=&quot;2017-12-23-managed-kubernetes-on-azure.md&quot;&gt;Managed Kubernetes on Microsoft Azure (English)&lt;/a&gt; for information on how to get up and running with Kubernetes on Azure.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;I have no affiliation with Microsoft Azure except using them through work.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;corrections&quot;&gt;Corrections&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;February 2020&lt;/strong&gt;: Some of my previous knowledge and assumptions were not correct when applied to a cloud + Docker environment, as &lt;a href=&quot;https://github.com/jnoller/kubernaughty/issues/46&quot;&gt;explained by
AKS PM Jesse Noller on GitHub&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;One of the issues is that even accessing a “data disk” will incur IOPS on the OS disk, and throttling of the OS disk will also constraint IOPS on the data disks.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Background&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;background&quot;&gt;Background&lt;/h2&gt;
&lt;p&gt;I’m part of a team at Equinor building an internal PaaS based on Kubernetes running on AKS (Azure managed Kubernetes). We use Prometheus for monitoring each cluster as well as InfluxDB for collecting metrics from k6io which runs continous tests on our public endpoints.&lt;/p&gt;
&lt;p&gt;A couple of weeks ago we discovered some potential problems with both Prometheus and InfluxDB with memory usage and restarts. High CPU usage of type &lt;code&gt;iowait&lt;/code&gt; suggested that there might be some disk issues contributing to the problems.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;iowait: “Percentage of time that the CPU or CPUs were idle during which the system had an outstanding disk I/O request.” (&lt;a href=&quot;https://support.hpe.com/hpsc/doc/public/display?docId=c02783994&quot;&gt;hpe.com&lt;/a&gt;). You can see &lt;code&gt;iowait&lt;/code&gt; on your Linux system by running &lt;code&gt;top&lt;/code&gt; and looking at the &lt;code&gt;wa&lt;/code&gt; percentage.&lt;/p&gt;
&lt;p&gt;PS: You can have a disk IO bottleneck even with low &lt;code&gt;iowait&lt;/code&gt;, and a high &lt;code&gt;iowait&lt;/code&gt; does not always indicate a disk IO bottleneck (&lt;a href=&quot;https://www.ibm.com/developerworks/community/blogs/AIXDownUnder/entry/iowait_a_misleading_indicator_of_i_o_performance54?lang=en&quot;&gt;ibm.com&lt;/a&gt;).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;First off we need to benchmark the underlying disk to get an understanding of it’s performance limits and characteristics. That is what we will cover in this post.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;MetricsMethodologies&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;metric-methodologies&quot;&gt;Metric Methodologies&lt;/h3&gt;
&lt;p&gt;There are two helpful methodologies when monitoring information systems. The first one is Utilization, Saturation and Errors (USE) from &lt;a href=&quot;http://www.brendangregg.com/usemethod.html&quot;&gt;Brendan Gregg&lt;/a&gt; and the second one is Rate, Errors, Duration (RED) from &lt;a href=&quot;https://www.slideshare.net/weaveworks/monitoring-microservices&quot;&gt;Tom Wilkie&lt;/a&gt;. RED is best suited when observing workloads and transactions while USE is best suited for observing resources.&lt;/p&gt;
&lt;p&gt;I’ll be using the USE method here. USE can be summarised as:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;For every resource, check utilization, saturation, and errors.&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;resource&lt;/strong&gt;: all physical server functional components (CPUs, disks, busses, …)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;utilization&lt;/strong&gt;: the average time that the resource was busy servicing work&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;saturation&lt;/strong&gt;: the degree to which the resource has extra work which it can’t service, often queued&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;errors&lt;/strong&gt;: the count of error events&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a id=&quot;StorageBackground&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;storage-background&quot;&gt;Storage Background&lt;/h3&gt;
&lt;p&gt;Disk usage has two dimensions, throughput/bandwidth(BW) and operations per second (IOPS), and the underlying storage system will have upper limits of how much data it can receive (BW) and the number of operations it can perform per second (IOPS).&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Background - harddrive types&lt;/strong&gt;: harddrives come in two types, Solid State Disks (SSD) and spindle (HDD). A SSD disk is a microship capable of permanently storing data while a HDD uses spinning platters to store data. HDDs have a fixed rate of rotation (RPM), typically 5.400 and 7.200 RPM for lower cost drives for home use and higher cost 10.000 and 15.000 RPM drives for server use. Over the last 20 years of HDDs their storage density has increased, but the RPM has largely stayed the same. A disk with twice the density (500GB to 1TB for example) can read twice as much data on a single rotation and thus increase the bandwidth significantly. However, reading or writing a random block still requires waiting for the disk to spin enough to reach the relevant sector on the disk. So IOPS has not increased much for HDDs and is still a low 125-150 IOPS for a 10.000 RPM enterprise disk. A SSD does not have any moving parts so is able to reach MUCH higher IOPS. A low end Samsung 960 EVO with 500GB capacity costs $150 and can achieve a whopping 330.000 IOPS! (&lt;a href=&quot;https://en.wikipedia.org/wiki/IOPS&quot;&gt;wikipedia.com&lt;/a&gt;)&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Background - access patterns&lt;/strong&gt;: The way a program uses storage also has a huge impact on the performance one can achieve. Sequential access is when we read or write a large file. When this happens the operating system and harddrive can optimize and “merge” operations so that we can read or write a much bigger chunk of data at a time. If we can read 1MB at a time 150 times per second we get 150MB/s of bandwidth. However, fully random access where the smallest chunk we read or write is a 4KB block the same 150 IOPS would only give a bandwidth of 0.6MB/s!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Background - cloud vs physical&lt;/strong&gt;: Now we know what HDDs are limited to a low IOPS and low IOPS combined with a random access pattern gives us a low overall bandwidth. There is a huge gotcha here when it comes to cloud. On Azure when using Premium Managed SSD the IOPS you are given is a factor of the disk size you provision (&lt;a href=&quot;https://azure.microsoft.com/en-us/pricing/details/managed-disks/&quot;&gt;microsoft.com&lt;/a&gt;). A 512GB disk is limited to 2.300 IOPS and 150MB/s. With 100% random access that only gives about 9MB/s of bandwidth!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Background - OS caching&lt;/strong&gt;: To overcome some of the limitations of the underlying disk (mostly IOPS) there are potentially several layers of caching involved. Linux file systems can have &lt;code&gt;writeback&lt;/code&gt; enabled which causes Linux to temporarily store data that is going to be written to disk in memory. This can give a big performance increase when there are sudden spikes of writes exceeding the performance of the underlying disk. It also increases the chance that operations can be &lt;code&gt;merged&lt;/code&gt; where several write operations to areas of the disk that are nearby can be executed as one. This caching works best for sudden peaks and will not necessarily be enough if there is continous random writes to disk. This caching also means that even though an application thinks it has saved some data to disk it can be lost in the case of a power outage or other failure. Applications can also explicitly request &lt;code&gt;direct&lt;/code&gt; access where every operation is persisted to disk before receiving a confirmation. This is a trade-off between performance and durability that needs to be decided based on the application itself and the environment.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Background - Azure caching&lt;/strong&gt;: Azure also provides read and write cache for its &lt;code&gt;disks&lt;/code&gt; which is enabled by default. As we will see soon for our use case it’s not a good idea to use.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;WhatToMeasure&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;what-to-measure&quot;&gt;What to measure?&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;These metrics are collected by the Prometheus &lt;code&gt;node-exporter&lt;/code&gt; and follows it’s naming. I’ve also created a dashboard that is available on &lt;a href=&quot;https://grafana.com/dashboards/9852&quot;&gt;Grafana.com&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;With the USE methodology as a guideline and the two separate but related “resources”, bandwidth and IOPS we can look for some useful metrics.&lt;/p&gt;
&lt;p&gt;Utilization:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;rate(node_disk_written_bytes_total)&lt;/code&gt; - Write bandwidth. The maximum is given by Azure and is 25MB/s for our disk size.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;rate(node_disk_writes_completed_total)&lt;/code&gt; - Write operations. The maximum is given by Azure and is 120 IOPS for our disk size.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;rate(node_disk_io_time_seconds_total)&lt;/code&gt; - Disk active time in percent. The time the disk was busy servicing requests. 100% means fully utilized.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Saturation:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;rate(node_cpu_seconds_total{mode=&quot;iowait&quot;}&lt;/code&gt; - CPU iowait. The percentage of time a CPU core is blocked from doing useful work because it’s waiting for an IO operation to complete (typically disk, but can also be network).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Useful calculated metrics:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;rate(node_disk_write_time_seconds_total) / rate(node_disk_writes_completed_total)&lt;/code&gt; - Write latency. How long from a write is requested until it’s completed.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;rate(node_disk_written_bytes_total) / rate(node_disk_writes_completed_total)&lt;/code&gt; - Write size. How big the &lt;strong&gt;average&lt;/strong&gt; write operation is. 4KB is minimum and indicates 100% random access while 512KB is maximum and indicates sequential access.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a id=&quot;HowToMeasureDisk&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;how-to-measure-disk&quot;&gt;How to measure disk&lt;/h2&gt;
&lt;p&gt;The best tool for measuring disk performance is &lt;code&gt;fio&lt;/code&gt;, even though it might seem a bit intimidating at first due to it’s insane number of options.&lt;/p&gt;
&lt;p&gt;Installing &lt;code&gt;fio&lt;/code&gt; on Ubuntu:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apt-get install fio&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;fio&lt;/code&gt; executes &lt;code&gt;jobs&lt;/code&gt; described in a file. Here is the top of our jobs file:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;[global]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;ioengine=libaio   # sync|libaio|mmap&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;group_reporting&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;thread&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;size=10g          # Size of test file&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cpus_allowed=1    # Only use this CPU core&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;runtime=300s      # Run test for 5 minutes&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;[test1]&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;filename=/tmp/fio-test-file&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;direct=1          # If value is true, use non-buffered I/O. Non-buffered I/O usually means O_DIRECT&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;readwrite=write   # read|write|randread|randwrite|readwrite|randrw&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;iodepth=1         # How many operations to queue to the disk&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;blocksize=4k&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The fields we will be changing for the various tests are &lt;code&gt;direct&lt;/code&gt;, &lt;code&gt;readwrite&lt;/code&gt;, &lt;code&gt;iodepth&lt;/code&gt; and &lt;code&gt;blocksize&lt;/code&gt;. Save the contents in a file named &lt;code&gt;jobs.fio&lt;/code&gt; and we run a test with &lt;code&gt;fio --sector test1 jobs.fio&lt;/code&gt; and wait until the test completes.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;PS: To run these tests on higher performance hardware and better caching you might want to set &lt;code&gt;runtime&lt;/code&gt; to &lt;code&gt;0&lt;/code&gt; to have the test run continously and monitor the metrics until performance reaches a steady-state.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;HowToMeasureDiskOnAKS&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;how-to-measure-disk-on-azure-kubernetes-service&quot;&gt;How to measure disk on Azure Kubernetes Service&lt;/h2&gt;
&lt;p&gt;For this testing we use a standard Prometheus installation collecting data from &lt;code&gt;node-exporter&lt;/code&gt; and visualizing data in Grafana. The dashboard I created for the testing can be found here: &lt;a href=&quot;https://grafana.com/dashboards/9852&quot;&gt;https://grafana.com/dashboards/9852&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;By default Kubernetes will schedule a Pod to any node that has enough memory and CPU for our workload. Since one of the tests we are going to run are on the OS disk we do not want the Pod to run on the same node as any other disk-intensive application, such as Prometheus.&lt;/p&gt;
&lt;p&gt;Look at which Pods are running with &lt;code&gt;kubectl get pods -o wide&lt;/code&gt; and look for a node that does not have any disk-intensive application.&lt;/p&gt;
&lt;p&gt;Then we tag that node with &lt;code&gt;kubectl label nodes aks-nodepool1-37707184-2 tag=disktest&lt;/code&gt;. This allows us later to specify that we want to run our testing Pod on that specific node.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;A StorageClass in Kubernetes is a specification of a underlying disk that Pods can request usage of through &lt;code&gt;volumeClaimTemplates&lt;/code&gt;. AKS comes with a default StorageClass &lt;code&gt;managed-premium&lt;/code&gt; that has caching enabled. Most of these tests require the Azure cache disabled so create a new StorageClass &lt;code&gt;managed-premium-retain-nocache&lt;/code&gt;:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kind: StorageClass&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apiVersion: storage.k8s.io/v1&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;metadata:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  name: managed-premium-retain-nocache&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;provisioner: kubernetes.io/azure-disk&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;reclaimPolicy: Retain&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;parameters:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  storageaccounttype: Premium_LRS&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  kind: Managed&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  cachingmode: None&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;You can add it to your cluster with:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl apply -f /attachments/2019-02-23-disk-performance-on-aks-part-1/storageclass.yaml&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;hr&gt;
&lt;p&gt;Next we create a StatefulSet that uses a &lt;code&gt;volumeClaimTemplate&lt;/code&gt; to request a 250GB Azure disk. This provisions a P15 Azure Premium SSD with 125MB/s bandwidth and 1100 IOPS:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl apply -f /attachments/2019-02-23-disk-performance-on-aks-part-1/ubuntu-statefulset.yaml&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Follow the progress of the Pod creation with &lt;code&gt;kubectl get pods -w&lt;/code&gt; and wait until it is &lt;code&gt;Running&lt;/code&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;When the Pod is &lt;code&gt;Running&lt;/code&gt; we can start a shell on it with &lt;code&gt;kubectl exec -it disk-test-0 bash&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;Once inside &lt;code&gt;bash&lt;/code&gt; on the Pod, we install &lt;code&gt;fio&lt;/code&gt;:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apt-get update &amp;amp;&amp;amp; apt-get install -y fio wget&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And save the contents of in the Pod:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;wget /attachments/2019-02-23-disk-performance-on-aks-part-1/jobs.fio&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Now we can run the different test sections one by one. &lt;strong&gt;PS: If you don’t specify a section &lt;code&gt;fio&lt;/code&gt; will run all the tests &lt;em&gt;simultaneously&lt;/em&gt;, which is not what we want.&lt;/strong&gt;&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;fio --section=test1 jobs.fio&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;fio --section=test2 jobs.fio&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;fio --section=test3 jobs.fio&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;fio --section=test4 jobs.fio&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;fio --section=test5 jobs.fio&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;fio --section=test6 jobs.fio&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;fio --section=test7 jobs.fio&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;fio --section=test8 jobs.fio&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;fio --section=test9 jobs.fio&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;Tests&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;test-results&quot;&gt;Test results&lt;/h2&gt;
&lt;p&gt;&lt;a id=&quot;Test1&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;test-1---learning-to-dislike-azure-cache&quot;&gt;Test 1 - Learning to dislike Azure Cache&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Sequential write, 4K block size, Azure Cache enabled, OS cache disabled. See &lt;a href=&quot;https://blog.stian.omg.lol/attachments/2019-02-23-disk-performance-on-aks-part-1/test1.md&quot;&gt;full fio test results&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;I run the first tests on the OS disk of a Kubernetes node. The OS disks have Azure caching enabled.&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;graph&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./test1.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;The first 1-2 minutes of the test I get very good performance of 45MB/s and ~11.500 IOPS but that drops to 0 very quickly as the cache is full and busy writing things to the underlying disk. When that happens everything freezes and I cannot even execute shell commands. After stopping the test the system still hangs for a bit while the cache empties.&lt;/p&gt;
&lt;p&gt;The maximum latency measured by &lt;code&gt;fio&lt;/code&gt; was 108751k usec. Or about 108 seconds!&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;For the first try of these tests a 20-30 second period of very fast writes (250MB/s) caused a 7-8 minutes hang while the cache emptied. Trying again caused another pattern of lower peak performance with shorter hangs in between. Very unpredictable.
I’m not sure what to make of this. It’s not acceptable that a Kubernetes node becomes unresponsive for many minutes following a short burst of writing. There are scattered recommendations online of disabling caching for write-heavy applications. Since I have not found any way to measure the Azure cache itself, the results are unpredictable and potentially very impactful as well as making it very hard to use the metrics we do have to evaluate application and storage behaviour I’ve concluded that it’s best to use data disks with caching disabled for our workloads (you cannot disable caching on an AKS node OS disk).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;Test2&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;test-2---disable-azure-cache---enable-os-cache&quot;&gt;Test 2 - Disable Azure Cache - enable OS cache&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Sequential write, 4K block size. &lt;strong&gt;Change: Azure cache disabled, OS caching enabled.&lt;/strong&gt; See &lt;a href=&quot;https://blog.stian.omg.lol/attachments/2019-02-23-disk-performance-on-aks-part-1/test2.md&quot;&gt;full fio test results&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;graph&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./test2.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;If we swap the Azure cache for the Linux OS cache we see that &lt;code&gt;iowait&lt;/code&gt; increases while the writing occurs. The application sees high write performance until the number of &lt;code&gt;Dirty bytes&lt;/code&gt; reaches a threshold of about 3.7GB of memory. The performance of the underlying disk is 125MB/s and 250 IOPS. Here we are throttled by the 125MB/s limit of the Azure P15 Premium SSD.&lt;/p&gt;
&lt;p&gt;Also notice that on sequential writes of 4K with OS caching the actual blocks written to disk is 512K which saves us a lot of IOPS. This will become important later.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Test3&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;test-3---disable-os-cache&quot;&gt;Test 3 - Disable OS cache&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Sequential write, 4K block size. &lt;strong&gt;Change: OS caching disabled.&lt;/strong&gt; See &lt;a href=&quot;https://blog.stian.omg.lol/attachments/2019-02-23-disk-performance-on-aks-part-1/test3.md&quot;&gt;full fio test results&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;graph&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./test3.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;By disabling the OS cache (&lt;code&gt;direct=1&lt;/code&gt;) the results are consistent and predictable. There is no &lt;code&gt;iowait&lt;/code&gt; since the application does not have multiple writes pending at the same time. Because of the 2-3ms latency of the disks we are not able to get more than about 400 IOPS. This gives us a meager 1.5MB/s even though the disk is limited to 1100 IOPS and 125MB/s. To reach that we need multiple simultaneous writes or a bigger IO depth (queue). &lt;code&gt;Disk active time&lt;/code&gt; is also 0% which indicates that the disk is not saturated.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;Test4&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;test-4---increase-io-depth&quot;&gt;Test 4 - Increase IO depth&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Sequential write, 4K block size, OS caching disabled. &lt;strong&gt;Change: IO depth 16.&lt;/strong&gt; See &lt;a href=&quot;https://blog.stian.omg.lol/attachments/2019-02-23-disk-performance-on-aks-part-1/test4.md&quot;&gt;full fio test results&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;graph&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./test4.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;For this test we only increase the IO depth from 1 to 16. IO depth is the number of write operations &lt;code&gt;fio&lt;/code&gt; will execute simultaneously. Since we are using &lt;code&gt;direct&lt;/code&gt; these will be queued by the OS for writing. We are now able to hit the performance limit of 1100 IOPS. &lt;code&gt;Disk active time&lt;/code&gt; is now steady at 100% indicating that we have saturated the disk.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;Test5&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;test-5---larger-block-size-smaller-io-depth&quot;&gt;Test 5 - Larger block size, smaller IO depth&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Sequential write, OS caching disabled. &lt;strong&gt;Change: 128K block size, IO depth 1.&lt;/strong&gt; See &lt;a href=&quot;https://blog.stian.omg.lol/attachments/2019-02-23-disk-performance-on-aks-part-1/test5.md&quot;&gt;full fio test results&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;graph&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./test5.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We increase the block size to 128KB and reduce the IO depth to 1 again. The write latency for larger blocks increase to ~5ms which gives us 200 IOPS and 28MB/s. The disk is not saturated.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;Test6&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;test-6---enable-os-cache&quot;&gt;Test 6 - Enable OS cache&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Sequential write, 256K block size, IO depth 1. &lt;strong&gt;Change: OS caching enabled.&lt;/strong&gt; See &lt;a href=&quot;https://blog.stian.omg.lol/attachments/2019-02-23-disk-performance-on-aks-part-1/test6.md&quot;&gt;full fio test results&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;graph&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./test6.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We have now enabled the OS cache/buffer (&lt;code&gt;direct=0&lt;/code&gt;). We can see that the writes hitting the disk are now merged to 512KB blocks. We are hitting the 125MB/s limit with about 250 IOPS. Enabling the cache also has other effects: CPU suddenly shows significant IO wait. The write latency shoots through the roof. Also note that the writing continued for 30-40 seconds after the test was done. &lt;strong&gt;This also means that the bandwidth and IOPS that &lt;code&gt;fio&lt;/code&gt; sees and reports is higher than what is actually hitting the disk.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;Test7&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;test-7---random-writes-small-block-size&quot;&gt;Test 7 - Random writes, small block size&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;IO depth 1, OS caching enabled. &lt;strong&gt;Change: Random write, 4K block size.&lt;/strong&gt; See &lt;a href=&quot;https://blog.stian.omg.lol/attachments/2019-02-23-disk-performance-on-aks-part-1/test7.md&quot;&gt;full fio test results&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;graph&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./test7.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Here we go from sequential writes to random writes. We are limited by IOPS. The average size of the blocks actually written to disks, and the IOPS required to hit the bandwidth limit is actually varying a bit throughout the test. The time taken to empty the cache is about as long as I ran the test (4-5 minutes).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;Test8&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;test-8---large-block-size&quot;&gt;Test 8 - Large block size&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Random write, OS caching enabled. &lt;strong&gt;Change: 256K block size, IO depth 16.&lt;/strong&gt; See &lt;a href=&quot;https://blog.stian.omg.lol/attachments/2019-02-23-disk-performance-on-aks-part-1/test8.md&quot;&gt;full fio test results&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;graph&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./test8.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Increasing the block size to 256K makes us bandwidth limited to 125MB/s.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;Conclusion&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Access patterns and block sizes have a tremendous impact on the amount of data we are able to write to disk.&lt;/p&gt;
</content:encoded><category>tech</category><category>kubernetes</category><category>azure</category></item><item><title>Managed Kubernetes on Microsoft Azure (English)</title><link>https://blog.stian.omg.lol/p/managed-kubernetes-on-microsoft-azure-english/</link><guid isPermaLink="true">https://blog.stian.omg.lol/p/managed-kubernetes-on-microsoft-azure-english/</guid><description>In this post we will set up a managed Kubernetes cluster from scratch using Azure CLI.</description><pubDate>Fri, 29 Dec 2017 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://blog.stian.omg.lol/_astro/2017-12-23-managed-kubernetes-on-azure.52ul0NBw_Z1AMWFr.jpg&quot; width=&quot;1600&quot; height=&quot;914&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;In this post we will set up a managed Kubernetes cluster from scratch using Azure CLI.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;A few days ago I wrote a walkthrough of &lt;a href=&quot;https://blog.stian.omg.lol/p/managed-kubernetes-p%C3%A5-microsoft-azure-norwegian/&quot;&gt;setting up Azure Container Service (AKS) in Norwegian&lt;/a&gt;. Someone asked me for an English version of that, and here it is.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Kubernetes(K8s) is becoming the de-facto standard for deploying container-based applications and workloads. Microsoft is currently in preview of their managed Kubernetes offering (Azure Kubernetes Service, AKS) which makes it easy to create a Kubernetes cluster and deploy workloads without the skill and time required to manage day-to-day operations of a Kubernetes-cluster, which today can be complex and time consuming.&lt;/p&gt;
&lt;p&gt;In this post we will set up a Kubernetes cluster from scratch using Azure CLI.
Table of contents&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#Background&quot;&gt;Background&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#Dockercontainers&quot;&gt;Docker containers&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Containerorchestration&quot;&gt;Container orchestration&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#QuickstartAKS&quot;&gt;Getting started with Azure Kubernetes - AKS&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#Caveats&quot;&gt;Caveats&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Preparations&quot;&gt;Preparations&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#AzureLogin&quot;&gt;Azure login&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#ActivateContainerService&quot;&gt;Activate ContainerService&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#CreateResourceGroup&quot;&gt;Create a resource group&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#CreateK8sCluster&quot;&gt;Create a Kubernetes cluster&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#InstallKubectl&quot;&gt;Install kubectl&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#InspectCluster&quot;&gt;Inspect cluster&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#StartNginx&quot;&gt;Start some nginx containere&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#NginxService&quot;&gt;Making nginx available with a service&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#ScaleCluster&quot;&gt;Scale cluster&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#DeleteCluster&quot;&gt;Delete cluster&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Bonusmaterial&quot;&gt;Bonus material&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#HelmIntro&quot;&gt;Deploying services with Helm&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#HelmMinecraft&quot;&gt;Deploy MineCraft with Helm&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#KubernetesDashboard&quot;&gt;Kubernetes Dashboard&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Conclusion&quot;&gt;Conclusion&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id=&quot;microsoft-azure&quot;&gt;Microsoft Azure&lt;/h4&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;a href=&quot;https://azure.microsoft.com/en-us/free/&quot;&gt;If you don’t have a Azure subscription already you can try services for $200 for 30 days.&lt;/a&gt; The VM size &lt;strong&gt;Standard_B2s&lt;/strong&gt; is Burstable, has 2vCPU, 4GB RAM, 8GB temp storage and costs roughly $38 / month. For $200 you can have a cluster of 3-4 B2s nodes plus traffic, loadbalancers and other additional costs.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;We have no affiliation with Microsoft Azure except their sponsorship of our startup &lt;a href=&quot;http://www.datadynamics.no/&quot;&gt;DataDynamics&lt;/a&gt; with cloud services for 24 months in their &lt;a href=&quot;https://bizspark.microsoft.com/&quot;&gt;BizSpark program&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;Background&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;background&quot;&gt;Background&lt;/h2&gt;
&lt;p&gt;&lt;a id=&quot;Dockercontainers&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;docker-containers&quot;&gt;Docker containers&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;We will not do a deep dive on Docker containers in this post, but here is a summary for those who are not familiar with it.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Docker is a way to package software so that it can run on the most popular platforms without worrying about installation, dependencies and to a certain degree, configuration.&lt;/p&gt;
&lt;p&gt;In addition, a Docker container uses the operating system of the host machine when it runs. Because of this it’s possible to run many more containers on the same host machine compared to running virtual machines.&lt;/p&gt;
&lt;p&gt;Here is a incomplete and rough comparison between a Docker container and a virtual machine:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Virtual machine&lt;/th&gt;
&lt;th&gt;Docker container&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Image size&lt;/td&gt;
&lt;td&gt;from 200MB to many GB&lt;/td&gt;
&lt;td&gt;from 10MB to 3-400MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Startup time&lt;/td&gt;
&lt;td&gt;60 seconds +&lt;/td&gt;
&lt;td&gt;1-10 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Memory usage&lt;/td&gt;
&lt;td&gt;256MB-512MB-1GB +&lt;/td&gt;
&lt;td&gt;2MB +&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security&lt;/td&gt;
&lt;td&gt;Good isolation between VMs&lt;/td&gt;
&lt;td&gt;Not as good isolation between containers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Building image&lt;/td&gt;
&lt;td&gt;Minutes&lt;/td&gt;
&lt;td&gt;Seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;PS&lt;/strong&gt; The numbers for virtual machines is taken from memory. I tried starting a MySQL virtual appliance on my laptop but VMware Player refuses to run because of Windows Hyper-V incompatibility. VMware Workstation refuses to run because of license issues and Oracle VirtualBox repeatedly gives me a nasty bluescreen. Hooray!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Protip&lt;/strong&gt; The smallest and fastest Docker images are built on Alpine Linux. For the webserver Nginx the Alpine-based image is 15MB compared to 108MB for the normal Debian-based image. PostgreSQL:Alpine is 38MB compared to 287MB with “full” OS. Last version of MySQL is 343MB but will in version 8 support Alpine Linux as well.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;To recap, some of the advantages of Docker containers are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Compatibility across platforms, Linux, Windows, MacOS.&lt;/li&gt;
&lt;li&gt;10-100x smaller size. Faster to download, build and upload.&lt;/li&gt;
&lt;li&gt;Memory usage only for application and not base OS.
&lt;ul&gt;
&lt;li&gt;Advantage when developing. Ability to run 10-20-30 containers on a development laptop.&lt;/li&gt;
&lt;li&gt;Advantage in production. Can reduce hardware/cloud costs considerably.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Near instant startup. Makes dynamic scaling of applications easier.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href=&quot;https://store.docker.com/editions/community/docker-ce-desktop-windows&quot;&gt;Download Docker for Windows here.&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;To start a MySQL database container from Windows CMD or Powershell:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;docker run --name mysql -p 3306:3306 -e MYSQL_RANDOM_ROOT_PASSWORD=true mysql&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Stop the container with:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;docker kill mysql&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;You can search for already built Docker images on &lt;a href=&quot;https://hub.docker.com/&quot;&gt;Docker Hub&lt;/a&gt;. It’s also possible to create private Docker repositories for your own software that you don’t want to be publicly available.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Containerorchestration&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;container-orchestration&quot;&gt;Container orchestration&lt;/h4&gt;
&lt;p&gt;Now that Docker container images has become the preferred way to package and distribute software on the Linux platform, there has emerged a need for systems to coordinate running and deploying these containers. Similar to the ecosystem of products VMware has built up around development and operation of virtual machines.&lt;/p&gt;
&lt;p&gt;Container orchestration systems have the responsibility for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Load balancing.&lt;/li&gt;
&lt;li&gt;Service discovery.&lt;/li&gt;
&lt;li&gt;Health checks.&lt;/li&gt;
&lt;li&gt;Automatic scaling and restarting of host nodes and containers.&lt;/li&gt;
&lt;li&gt;Zero downtime upgrades (rolling deploys).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Until recently the ecosystem around container orchestration has been fragmented, and the most popular alternatives have been:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://kubernetes.io/&quot;&gt;Kubernetes&lt;/a&gt; (Originaly from Google, now managed by CNCF, the Cloud Native Computing Foundation)&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.docker.com/engine/swarm/&quot;&gt;Swarm&lt;/a&gt; (From the maker of Docker)&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;http://mesos.apache.org/&quot;&gt;Mesos&lt;/a&gt; (From Apache Software Foundation)&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/coreos/fleet&quot;&gt;Fleet&lt;/a&gt; (From CoreOS)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But the last year there has been a convergence towards Kubernetes as the preferred solution.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;7 February
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://coreos.com/blog/migrating-from-fleet-to-kubernetes.html&quot;&gt;CoreOS announces that they are removing Fleet from Container Linux and recommends Kubernetes&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;27 July
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://azure.microsoft.com/en-us/blog/announcing-cncf/&quot;&gt;Microsoft joins the CNCF&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;9 August
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://techcrunch.com/2017/08/09/aws-joins-the-cloud-native-computing-foundation/&quot;&gt;Amazon Web Services join the CNCF&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;29 August
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.geekwire.com/2017/now-vmware-pivotal-cncf-becoming-hub-enterprise-tech/&quot;&gt;VMware and Pivotal joins the CNCF&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;17 September
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://techcrunch.com/2017/09/13/oracle-joins-the-cloud-native-computing-foundation-as-a-platinum-member/&quot;&gt;Oracle joins the CNCF&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;17 October
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.theregister.co.uk/2017/10/17/docker_ee_kubernetes_support/&quot;&gt;Docker announces native support for Kubernetes in addition to it’s own Swarm product&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;24 October
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://azure.microsoft.com/en-us/blog/introducing-azure-container-service-aks-managed-kubernetes-and-azure-container-registry-geo-replication/&quot;&gt;Microsoft Azure announces the managed Kubernetes service AKS&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;29 November
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://aws.amazon.com/blogs/aws/amazon-elastic-container-service-for-kubernetes/&quot;&gt;Amazon Web Services announces the managed Kubernetes service EKS&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Especially the last two news items are important. Deploying and running your own Kubernetes-installation requires time and skills (&lt;a href=&quot;https://stripe.com/blog/operating-kubernetes&quot;&gt;Read how Stripe used 5 months to trust running Kubernetes in production, just for batch jobs.&lt;/a&gt;)&lt;/p&gt;
&lt;p&gt;Until now the choice has been running your own Kubernetes cluster or using Google Container Engine which has been &lt;a href=&quot;https://cloudplatform.googleblog.com/2014/11/unleashing-containers-and-kubernetes-with-google-compute-engine.html&quot;&gt;using Kubernetes since 2014&lt;/a&gt;. Many of us feel a certain discomfort by locking ourselves to one provider. But this is now changing when you can develop infrastructure on Kubernetes and choose between the 3 large cloud providers in addition to running your own cluster if wanted. &lt;strong&gt;*&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;*&lt;/strong&gt; Kubernetes is a fast moving project, and features might be available on the different platforms on different timelines.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;QuickstartAKS&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;getting-started-with-azure-kubernetes---aks&quot;&gt;Getting started with Azure Kubernetes - AKS&lt;/h2&gt;
&lt;p&gt;&lt;a id=&quot;Caveats&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;caveats&quot;&gt;Caveats&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;This guide is based on the documentation on &lt;a href=&quot;https://docs.microsoft.com/en-us/azure/aks/kubernetes-walkthrough&quot;&gt;Microsoft.com&lt;/a&gt;. Setting up a Azure Kubernetes cluster did not work in the beginning of December, but today, 23. December, it seems to work fairly well. But, upgrading the cluster from Kubernetes 1.7 to 1.8 for example does NOT work.&lt;/p&gt;
&lt;p&gt;AKS is in Preview and Azure are working continuously to make AKS stable and to support as many Kubernetes-features as possible. Amazon Web Services has a similar closed invite-only Preview currently while working on stability and features.&lt;/p&gt;
&lt;p&gt;Both Azure and AWS expresses expectations about their Kubernetes offerings will be ready for production in 2018.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;Preparations&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;preparations&quot;&gt;Preparations&lt;/h3&gt;
&lt;p&gt;You need Azure-CLI (version 2.0.21 or newer) to execute the &lt;code&gt;az&lt;/code&gt; commands:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://aka.ms/InstallAzureCliWindows&quot;&gt;Download Azure-CLI here&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.microsoft.com/en-us/cli/azure/install-azure-cli?view=azure-cli-latest&quot;&gt;Information about Azure-CLI on MacOS and Linux here&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;All commands executed in Windows PowerShell.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;AzureLogin&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;azure-login&quot;&gt;Azure login&lt;/h3&gt;
&lt;p&gt;Log on to Azure:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az login&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;You will get a link to open in your browser together with an authentication code. Enter the code on the webpage and &lt;code&gt;az login&lt;/code&gt; will save the login information so that you will not have to authenticate again on the same machine.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;PS&lt;/strong&gt; The login information gets saved in &lt;code&gt;C:\Users\Username\.azure\&lt;/code&gt;. You have to make sure nobody can access these files. They will then have full access to your Azure account.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;ActivateContainerService&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;activate-containerservice&quot;&gt;Activate ContainerService&lt;/h3&gt;
&lt;p&gt;Since AKS is in Preview/Beta, you explicitly have to activate it in your subscription to get access to the &lt;code&gt;aks&lt;/code&gt; subcommands.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az provider register -n Microsoft.ContainerService&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az provider show -n Microsoft.ContainerService&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;CreateResourceGroup&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;create-a-resource-group&quot;&gt;Create a resource group&lt;/h3&gt;
&lt;p&gt;Here we create a resource group named “my_aks_rg” in Azure region West Europe.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az group create --name my_aks_rg --location westeurope&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Protip&lt;/strong&gt;
To see a list of all available Azure regions, use the command &lt;code&gt;az account list-locations --output table&lt;/code&gt;. &lt;strong&gt;PS&lt;/strong&gt; AKS might not be available in all regions yet!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;CreateK8sCluster&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;create-kubernetes-cluster&quot;&gt;Create Kubernetes cluster&lt;/h3&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az aks create --resource-group my_aks_rg --name my_cluster --node-count 3 --generate-ssh-keys --node-vm-size Standard_B2s --node-osdisk-size 128 --kubernetes-version 1.8.2&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;--node-count&lt;/code&gt;
&lt;ul&gt;
&lt;li&gt;Number of agent(host) nodes available to run containers&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--generate-ssh-keys&lt;/code&gt;
&lt;ul&gt;
&lt;li&gt;Creates and prints a SSH key which can be used for SSHing directly to the agent nodes.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--node-vm-size&lt;/code&gt;
&lt;ul&gt;
&lt;li&gt;Which size Azure VMs the agent nodes should be created as. To see available sizes use &lt;code&gt;az vm list-sizes -l westeurope --output table&lt;/code&gt; and &lt;a href=&quot;https://docs.microsoft.com/en-us/azure/virtual-machines/linux/sizes&quot;&gt;Microsofts webpages&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--node-osdisk-size&lt;/code&gt;
&lt;ul&gt;
&lt;li&gt;Disk size of the agent nodes in GB. &lt;strong&gt;PS&lt;/strong&gt; Containers can be stopped and moved to another host if Kubernetes finds it necessary or if a agent node disappears. All data saved locally in the container will be gone. If saving data permanently use Kubernetes PersistentVolumes and not the local agent node or container disks.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--kubernetes-version&lt;/code&gt;
&lt;ul&gt;
&lt;li&gt;Which Kubernetes version to install. Azure does NOT necessarily install the last version by default, and currently upgrading with &lt;code&gt;az aks upgrade&lt;/code&gt; does not work. Latest version available right now is 1.8.2. It’s recommended to use the latest available version since there is a lot of changes from version to version. The documentation is also much better for newer versions.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Save the output of the command in a file in a secure location. It contains keys that can be used to connect to the cluster with SSH. Even though that should not in theory be necessary.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;InstallKubectl&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;install-kubectl&quot;&gt;Install kubectl&lt;/h3&gt;
&lt;p&gt;&lt;code&gt;kubectl&lt;/code&gt; is the client which performs all operations against your Kubernetes cluster. Azure CLI can install &lt;code&gt;kubectl&lt;/code&gt; for you:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az aks install-cli&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;After &lt;code&gt;kubectl&lt;/code&gt; is installed we need to get login information so that &lt;code&gt;kubectl&lt;/code&gt; can communicate with the Kubernetes cluster.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az aks get-credentials --resource-group my_aks_rg --name my_cluster&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The login information is saved in &lt;code&gt;C:\Users\Username\.kube\config&lt;/code&gt;. Keep these files secure as well.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Protip&lt;/strong&gt; When you have several Kubernetes clusters you can change which one &lt;code&gt;kubectl&lt;/code&gt; talks to with &lt;code&gt;kubectl config get-contexts&lt;/code&gt; and &lt;code&gt;kubectl config set-context my_cluster&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;InspectCluster&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;inspect-cluster&quot;&gt;Inspect cluster&lt;/h3&gt;
&lt;p&gt;To check that the cluster and &lt;code&gt;kubectl&lt;/code&gt; works we start with a couple of commands.&lt;/p&gt;
&lt;p&gt;See all agent nodes and status:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;gt; kubectl get nodes&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAME                       STATUS    AGE       VERSION&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;aks-nodepool1-16970026-0   Ready     15m       v1.8.2&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;aks-nodepool1-16970026-1   Ready     15m       v1.8.2&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;aks-nodepool1-16970026-2   Ready     15m       v1.8.2&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;See all services, pods and deployments:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;gt; kubectl get all --all-namespaces&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAMESPACE     NAME                                          READY     STATUS    RESTARTS   AGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kube-system   po/kubernetes-dashboard-6fc8cf9586-frpkn      1/1       Running   0          3d&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAMESPACE     NAME                          CLUSTER-IP     EXTERNAL-IP     PORT(S)           AGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kube-system   svc/kubernetes-dashboard      10.0.161.132   &amp;lt;none&amp;gt;          80/TCP            3d&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAMESPACE     NAME                             DESIRED   CURRENT   UP-TO-DATE   AVAILABLE   AGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kube-system   deploy/kubernetes-dashboard      1         1         1            1           3d&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAMESPACE     NAME                                    DESIRED   CURRENT   READY     AGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kube-system   rs/kubernetes-dashboard-6fc8cf9586      1         1         1         3d&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This is just some of the output from this command. You do not have to know what the resources in the &lt;code&gt;kube-system&lt;/code&gt; namespace does. That is part of the intention when Microsoft is managing our cluster for us.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Namespaces&lt;/strong&gt;
In Kubernetes there is something called Namespaces. Resources in one namespace does not have automatic access to resources in another namespace. The services that runs Kubernetes itself use the namespace &lt;code&gt;kube-system&lt;/code&gt;. The &lt;code&gt;kubectl&lt;/code&gt; command by default only shows you resources in the &lt;code&gt;default&lt;/code&gt; namespace, unless you specify &lt;code&gt;--all-namespaces&lt;/code&gt; or &lt;code&gt;--namespace=xx&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;StartNginx&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;start-some-nginx-containers&quot;&gt;Start some nginx containers&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;An instance of a running container in Kubernetes is called a &lt;strong&gt;Pod&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;nginx&lt;/code&gt; is a fast and flexible web server.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Now that the clsuter is up we can start rolling out services and deployments on it.&lt;/p&gt;
&lt;p&gt;Lets start with creating a Deployment consiting of 3 containers all running the &lt;code&gt;nginx:mainline-alpine&lt;/code&gt; image from &lt;a href=&quot;https://hub.docker.com/r/_/nginx/&quot;&gt;Docker hub&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;nginx-dep.yaml&lt;/strong&gt; looks like this:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apiVersion: apps/v1beta2&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kind: Deployment&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;metadata:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  name: nginx-deployment&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  labels:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    app: nginx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;spec:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  replicas: 3&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  selector:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    matchLabels:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      app: nginx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  template:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    metadata:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      labels:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        app: nginx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    spec:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      containers:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      - name: nginx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        image: nginx:mainline-alpine&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        ports:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        - containerPort: 80&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Load this into the cluster with &lt;code&gt;kubectl create&lt;/code&gt;:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl create -f /attachments/2017-12-23-managed-kubernetes-on-azure/nginx-dep.yaml&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This command creates the resources described in the file. &lt;code&gt;kubectl&lt;/code&gt; can read files either from your local disk or from a web URL.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;After making changes to a resource definition (&lt;code&gt;.yaml&lt;/code&gt; file), you can update the resources in the cluster with &lt;code&gt;kubetl replace -f resource.yaml&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;We can verify that the Deployment is ready:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;gt; kubectl get deploy&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAME               DESIRED   CURRENT   UP-TO-DATE   AVAILABLE   AGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;nginx-deployment   3         3         3            3           10m&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;We can also get the actual Pods that are running:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;gt; kubectl get pods&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAME                                READY     STATUS    RESTARTS   AGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;nginx-deployment-569477d6d8-dqwx5   1/1       Running   0          10m&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;nginx-deployment-569477d6d8-xwzpw   1/1       Running   0          10m&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;nginx-deployment-569477d6d8-z5tfk   1/1       Running   0          10m&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Logger&lt;/strong&gt; We can view logs from one pod with &lt;code&gt;kubectl logs nginx-deployment-569477d6d8-xwzpw&lt;/code&gt;. But since we in this case don’t know which Pod ends up getting an incomming request we can view logs from all the Pods which have &lt;code&gt;app=nginx&lt;/code&gt; label: &lt;code&gt;kubectl logs -lapp=nginx&lt;/code&gt;. The use of &lt;code&gt;app=nginx&lt;/code&gt; is our choice in &lt;code&gt;nginx-dep.yaml&lt;/code&gt; when we configured &lt;code&gt;spec.template.metadata.labels: app: nginx&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;NginxService&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;making-nginx-available-with-a-service&quot;&gt;Making nginx available with a service&lt;/h3&gt;
&lt;p&gt;To send traffic to our new Pods we need to create a &lt;strong&gt;Service&lt;/strong&gt;. A service consists of one or more Pods which are chosen based on different criteria, for example which labels they have and whether the Pods are Running and Ready.&lt;/p&gt;
&lt;p&gt;Lets create a service which forwards traffic to all Pods with label &lt;code&gt;app: nginx&lt;/code&gt; and are listening to port 80. In addition we make the service available via a LoadBalancer:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;nginx-svc.yaml&lt;/strong&gt; looks like this:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apiVersion: v1&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kind: Service&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;metadata:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  name: nginx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  labels:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    app: nginx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;spec:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  type: LoadBalancer&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  ports:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  - port: 80&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    name: http&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    targetPort: 80&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  selector:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    app: nginx&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;We tell Kubernetes to create our service with &lt;code&gt;kubectl create&lt;/code&gt; as usual:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl create -f /attachments/2017-12-23-managed-kubernetes-on-azure/nginx-svc.yaml&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;We can then wait and see which IP-address Azure assigns our service:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;gt; kubectl get svc -w&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAME         CLUSTER-IP   EXTERNAL-IP     PORT(S)        AGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;nginx        10.0.24.11   13.95.173.255   80:31522/TCP   15m&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;PS&lt;/strong&gt; It can take a few minutes for Azure to allocate and assign a Public IP for us. In the mean time &lt;code&gt;&amp;lt;pending&amp;gt;&lt;/code&gt; will appear under EXTERNAL-IP.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;A simple &lt;strong&gt;Welcome to nginx&lt;/strong&gt; webpage should now be available on &lt;a href=&quot;http://13.95.173.255&quot;&gt;http://13.95.173.255&lt;/a&gt; (&lt;em&gt;remember to replace with your own External-IP&lt;/em&gt;).&lt;/p&gt;
&lt;p&gt;We can also delete the service and deployment afterwards:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl delete svc nginx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl delete deploy nginx-deployment&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;ScaleCluster&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;scaling-the-cluster&quot;&gt;Scaling the cluster&lt;/h3&gt;
&lt;p&gt;If we want to change the number of agent nodes running Pods we can do that via Azure-CLI:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az aks scale --name my_cluster --resource-group my_aks_rg --node-count 5&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;Currently all nodes will be created with the same size as when we created the cluster. AKS will probably get support for &lt;a href=&quot;https://cloud.google.com/kubernetes-engine/docs/concepts/node-pools&quot;&gt;&lt;strong&gt;node-pools&lt;/strong&gt;&lt;/a&gt; next year. That will allow for creating different groups of nodes with different size and operating systems, both Linux and Windows.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;DeleteCluster&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;delete-cluster&quot;&gt;Delete cluster&lt;/h3&gt;
&lt;p&gt;You can delete the whole cluster like this:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az aks delete --name my_cluster --resource-group my_aks_rg --yes&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;Bonusmaterial&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;bonus-material&quot;&gt;Bonus material&lt;/h2&gt;
&lt;p&gt;Here is some bonus material if you want to go a bit further with Kubernetes.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;HelmIntro&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;deploying-services-with-helm&quot;&gt;Deploying services with Helm&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://helm.sh/&quot;&gt;Helm&lt;/a&gt; is a package manager and library of software that is ready to be deployed on a Kubernetes cluster.&lt;/p&gt;
&lt;p&gt;Start by downloading the &lt;a href=&quot;https://github.com/kubernetes/helm/releases&quot;&gt;Helm-client&lt;/a&gt;. It will read login information etc. from the same location as &lt;code&gt;kubectl&lt;/code&gt; automatically.&lt;/p&gt;
&lt;p&gt;Install the Helm-server (&lt;strong&gt;Tiller&lt;/strong&gt;) on the Kubernetes cluster and update the package library:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;helm init&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;helm repo update&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;See available packages (&lt;strong&gt;Charts&lt;/strong&gt;) with &lt;code&gt;helm search&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;HelmMinecraft&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;deploy-minecraft-with-helm&quot;&gt;Deploy MineCraft with Helm&lt;/h4&gt;
&lt;p&gt;Lets deploy a MineCraft server installation on our cluster, just because we can :-)&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;helm install --name stians --set minecraftServer.eula=true stable/minecraft&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;--set&lt;/code&gt; overrides one or more of the standard values configured in the package. The MineCraft package is made in a way where it does not start without accepting the user license agreement by setting the variable &lt;code&gt;minecraftServer.eula&lt;/code&gt;. All the variables that can be set in the MineCraft package are &lt;a href=&quot;https://github.com/kubernetes/charts/blob/master/stable/minecraft/values.yaml&quot;&gt;documented here&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Then we wait for Azure to assign us a Public IP:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;gt; kubectl get svc -w&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;stians-minecraft   10.0.237.0   13.95.172.192   25565:30356/TCP   3m&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Now we can connect to our MineCraft server on &lt;code&gt;13.95.172.192:25565&lt;/code&gt;!&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Kubernetes in MineCraft on Kubernetes&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./minecraft-k8s.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;KubernetesDashboard&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;kubernetes-dashboard&quot;&gt;Kubernetes Dashboard&lt;/h3&gt;
&lt;p&gt;Kubernetes also has a graphic web user-interface which makes it a bit easier to see which resources are in the cluster, view logs and even open a remote shell inside a running Pod, among other things.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;gt; kubectl proxy&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Starting to serve on 127.0.0.1:8001&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;kubectl&lt;/code&gt; encrypts and tunnels the traffic to the Kubernetes API servers. The dashboard is available on &lt;a href=&quot;http://127.0.0.1:8001/ui/&quot;&gt;http://127.0.0.1:8001/ui/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Kubernetes Dashboard&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./k8s-dash.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Conclusion&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;I hope you enjoy Kubernetes as much as I have. The learning curve can be a bit steep in the beginning, but it does not take long before you are productive.&lt;/p&gt;
&lt;p&gt;Look at the &lt;a href=&quot;https://v1-8.docs.kubernetes.io/docs/tutorials/&quot;&gt;official guides on Kubernetes.io&lt;/a&gt; to learn more about defining different types of resources and services to run on Kubernetes. &lt;strong&gt;PS: There are big changes from version to version so make sure you use the documentation for the correct version!&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Kubernetes also have a very active Slack-community on &lt;a href=&quot;http://slack.k8s.io/&quot;&gt;kubernetes.slack.com&lt;/a&gt; that is worthwhile to check out.&lt;/p&gt;
</content:encoded><category>tech</category><category>kubernetes</category><category>azure</category></item><item><title>Managed Kubernetes på Microsoft Azure (Norwegian)</title><link>https://blog.stian.omg.lol/p/managed-kubernetes-p%C3%A5-microsoft-azure-norwegian/</link><guid isPermaLink="true">https://blog.stian.omg.lol/p/managed-kubernetes-p%C3%A5-microsoft-azure-norwegian/</guid><description>I denne posten setter vi opp et managed Kubernetes cluster fra scratch ved bruk av Azure CLI.</description><pubDate>Mon, 25 Dec 2017 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://blog.stian.omg.lol/_astro/2017-12-23-managed-kubernetes-on-azure.52ul0NBw_Z1AMWFr.jpg&quot; width=&quot;1600&quot; height=&quot;914&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;I denne posten setter vi opp et managed Kubernetes cluster fra scratch ved bruk av Azure CLI.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Update 29. Dec: There is an &lt;a href=&quot;https://blog.stian.omg.lol/p/managed-kubernetes-on-microsoft-azure-english/&quot;&gt;English version of this post here.&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Kubernetes (K8s) er i ferd med å bli de-facto standard for deployments av kontainer-baserte applikasjoner. Microsoft har nå preview av deres managed Kubernetes tjeneste (Azure Kubernetes Service, AKS) som gjør det enkelt å opprette et Kubernetes cluster og rulle ut tjenester uten å måtte ha kompetanse og tid til den daglige driften av selve Kubernetes-clusteret, som per i dag kan være relativt komplisert og tidkrevende.&lt;/p&gt;
&lt;p&gt;I denne posten setter vi opp et Kubernetes cluster fra scratch ved bruk av Azure CLI.
Table of contents&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#Bakgrunn&quot;&gt;Bakgrunn&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#Dockercontainers&quot;&gt;Docker containers&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Containerorchestration&quot;&gt;Container orchestration&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#OppretteAKS&quot;&gt;Kom i gang med Azure Kubernetes - AKS&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#Forbehold&quot;&gt;Forbehold&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Forberedelser&quot;&gt;Forberedelser&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Azureinnlogging&quot;&gt;Azure innlogging&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#AktiverContainerService&quot;&gt;Aktiver ContainerService&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#OpprettResourceGroup&quot;&gt;Opprett en resource group&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#OpprettK8sCluster&quot;&gt;Opprette Kubernetes cluster&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#InstallerKubectl&quot;&gt;Installer kubectl&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#InspiserCluster&quot;&gt;Inspiser cluster&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#StarteNginx&quot;&gt;Starte noen nginx containere&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#NginxService&quot;&gt;Gjøre nginx tilgjengelig med en tjeneste&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#ScaleCluster&quot;&gt;Skalere cluster&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#DeleteCluster&quot;&gt;Slette cluster&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Bonusmateriale&quot;&gt;Bonusmateriale&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#HelmIntro&quot;&gt;Rulle ut tjenester med Helm pakker&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#HelmMinecraft&quot;&gt;MineCraft server med Helm&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#KubernetesDashboard&quot;&gt;Kubernetes Dashboard&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Konklusjon&quot;&gt;Konklusjon&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id=&quot;microsoft-azure&quot;&gt;Microsoft Azure&lt;/h4&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;a href=&quot;https://azure.microsoft.com/en-us/free/&quot;&gt;Hvis du ikke har Azure fra før kan du prøve tjenester for $200 i 30 dager.&lt;/a&gt; VM typen &lt;strong&gt;Standard_B2s&lt;/strong&gt; er Burstable, har 2vCPU, 4GB RAM, 8GB temp storage og koster ~$38 / mnd. For $200 kan du ha et cluster på 3-4 B2s noder plus trafikkostnad, lastbalanserere og andre nødvendige tjenester.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Vi har ingen tilknytning til Microsoft bortsett fra at de sponser vår startup &lt;a href=&quot;http://www.datadynamics.no/&quot;&gt;DataDynamics&lt;/a&gt; med cloud-tjenester i 24 mnd i deres &lt;a href=&quot;https://bizspark.microsoft.com/&quot;&gt;BizSpark program&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;Bakgrunn&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;bakgrunn&quot;&gt;Bakgrunn&lt;/h2&gt;
&lt;p&gt;&lt;a id=&quot;Dockercontainers&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;docker-containers&quot;&gt;Docker containers&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Vi tar ikke for oss Docker containers i dybden i denne posten, men her er en kort oppsummering for de som ikke er kjent med teknologien.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Docker er en måte å pakketere programvare slik at det kan kjøres på samtlige populære platformer uten å måtte bruke mye tid på dependencies, oppsett og konfigurasjon.&lt;/p&gt;
&lt;p&gt;I tillegg bruker en Docker container operativsystemet på vertsmaskinen når den kjører. Dette gjør at en kan kjøre mange flere containere på samme vertsmaskin sammenlignet med virtuelle maskiner.&lt;/p&gt;
&lt;p&gt;Her er en ufullstendig og grov sammenligning mellom en Docker container og en virtuell maskin:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Virtuel maskin&lt;/th&gt;
&lt;th&gt;Docker container&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Image størrelse&lt;/td&gt;
&lt;td&gt;fra 200MB til mange GB&lt;/td&gt;
&lt;td&gt;fra 10MB til 3-400MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Oppstartstid&lt;/td&gt;
&lt;td&gt;60 sekunder +&lt;/td&gt;
&lt;td&gt;1-10 sekunder&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Minnebruk&lt;/td&gt;
&lt;td&gt;256MB-512MB-1GB +&lt;/td&gt;
&lt;td&gt;2MB +&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sikkerhet&lt;/td&gt;
&lt;td&gt;God isolasjon mellom VM&lt;/td&gt;
&lt;td&gt;Dårligere isolasjon mellom containere&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bygge image&lt;/td&gt;
&lt;td&gt;Minutter&lt;/td&gt;
&lt;td&gt;Sekunder&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;PS&lt;/strong&gt; Tallene for virtuelle maskiner er tatt fra hukommelsen. Jeg forsøkte å starte en MySQL virtuell appliance på min laptop men VMware Player nekter å kjøre pga inkompatibilitet med Windows Hyper-V. VMware Workstation nekter å kjøre pga utgått lisens og Oracle VirtualBox gir en nasty bluescreen gang på gang. Hooray!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Protip&lt;/strong&gt; De minste og raskeste Docker imagene er bygget på Alpine Linux. For webserveren Nginx er det Alpine-baserte imaget 15MB mot det Debian-baserte imaget på 108MB. PostgreSQL:Alpine er 38MB mot 287MB. Siste versjon av MySQL er 343MB men vil i versjon 8 støtte Alpine Linux også.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Noen av fordelene med Docker containers er altså:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Kompatibilitet på tvers av platformer, Linux, Windows og MacOS.&lt;/li&gt;
&lt;li&gt;10-100x mindre størrelse. Raskere å laste ned, raskere å bygge, raskere å laste opp.&lt;/li&gt;
&lt;li&gt;Minnebruk kun for applikasjon og ikke eget OS.
&lt;ul&gt;
&lt;li&gt;Fordel under utvikling, kan kjøre 10-20-30 Docker containere samtidig på en laptop.&lt;/li&gt;
&lt;li&gt;Fordel i produksjon, kan redusere hardware utgifter betraktelig.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Oppstart på få sekunder. Gjør dynamisk skalering av applikasjoner mye enklere.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href=&quot;https://store.docker.com/editions/community/docker-ce-desktop-windows&quot;&gt;Last ned Docker for Windows her.&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Og start en MySQL database fra Windows CMD eller Powershell:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;docker run --name mysql -p 3306:3306 -e MYSQL_RANDOM_ROOT_PASSWORD=true mysql&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Stop containeren med:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;docker kill mysql&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;En kan søke etter ferdige Docker images på &lt;a href=&quot;https://hub.docker.com/&quot;&gt;Docker Hub&lt;/a&gt;. Det er også mulig å lage private Docker repositories for egen programvare som ikke skal være tilgjengelig for omverden.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Containerorchestration&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;container-orchestration&quot;&gt;Container orchestration&lt;/h4&gt;
&lt;p&gt;Etter som Docker containers har blitt den foretrukne måten å pakke og distribuere programvare på Linux platformen de siste par årene har det vokst frem et behov for systemer som kan samkjøre drift og utrulling av disse containerene. Ikke ulikt det økosystemet av produkter VMware har bygget opp rundt utvikling og drift av virtuelle maskiner.&lt;/p&gt;
&lt;p&gt;Container orchestration systemene har som oppgave å sørge for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Lastbalansering.&lt;/li&gt;
&lt;li&gt;Service discovery.&lt;/li&gt;
&lt;li&gt;Health checks.&lt;/li&gt;
&lt;li&gt;Automatisk skalering og restarting av vertsmaskiner og containere.&lt;/li&gt;
&lt;li&gt;Oppgraderinger uten nedetid (rolling deploy).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Frem til nylig har økosystemet rundt container orchestration vært fragmentert og de mest populære alternativene har vært:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://kubernetes.io/&quot;&gt;Kubernetes&lt;/a&gt; (Opprinnelig fra Google, nå styrt av CNCF, Cloud Native Computing Foundation)&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.docker.com/engine/swarm/&quot;&gt;Swarm&lt;/a&gt; (Fra produsenten bak Docker)&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;http://mesos.apache.org/&quot;&gt;Mesos&lt;/a&gt; (Fra Apache Software Foundation)&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/coreos/fleet&quot;&gt;Fleet&lt;/a&gt; (Fra CoreOS)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Men det siste året har det vært en konvergens mot Kubernetes som foretrukket løsning.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;7 februar
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://coreos.com/blog/migrating-from-fleet-to-kubernetes.html&quot;&gt;CoreOS annonserer at de fjerner Fleet fra Container Linux og anbefaler Kubernetes&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;27 juli
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://azure.microsoft.com/en-us/blog/announcing-cncf/&quot;&gt;Microsoft slutter seg til CNCF&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;9 august
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://techcrunch.com/2017/08/09/aws-joins-the-cloud-native-computing-foundation/&quot;&gt;Amazon Web Services slutter seg til CNCF&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;29 august
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.geekwire.com/2017/now-vmware-pivotal-cncf-becoming-hub-enterprise-tech/&quot;&gt;VMware og Pivotal slutter seg til CNCF&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;17 september
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://techcrunch.com/2017/09/13/oracle-joins-the-cloud-native-computing-foundation-as-a-platinum-member/&quot;&gt;Oracle slutter seg til CNCF&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;17 oktober
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://www.theregister.co.uk/2017/10/17/docker_ee_kubernetes_support/&quot;&gt;Docker annonserer native støtte for Kubernetes i tillegg til sitt eget Swarm produkt&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;24 oktober
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://azure.microsoft.com/en-us/blog/introducing-azure-container-service-aks-managed-kubernetes-and-azure-container-registry-geo-replication/&quot;&gt;Microsoft Azure annonserer managed Kubernetes med tjenesten AKS&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;29 november
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://aws.amazon.com/blogs/aws/amazon-elastic-container-service-for-kubernetes/&quot;&gt;Amazon Web Services annonserer managed Kubernetes med tjenesten EKS&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;De to siste nyhetene er spesielt viktige. Å drifte sin egen Kubernetes-installasjon krever tid og kompetanse. (&lt;a href=&quot;https://stripe.com/blog/operating-kubernetes&quot;&gt;Les hvordan Stripe brukte 5 måneder på å bli fortrolig med å drifte sitt eget Kubernetes cluster, bare for batch jobs.&lt;/a&gt;)&lt;/p&gt;
&lt;p&gt;Frem til nå har valget vært mellom å drifte sitt eget Kubernetes cluster eller bruke Google Container Engine som har &lt;a href=&quot;https://cloudplatform.googleblog.com/2014/11/unleashing-containers-and-kubernetes-with-google-compute-engine.html&quot;&gt;brukt Kubernetes siden 2014&lt;/a&gt;. Mange av oss føler et visst ubehag ved å låse oss til én tilbyder. Men dette er nå anderledes når en kan utvikle infrastruktur på Kubernetes, og velge tilnærmet fritt &lt;strong&gt;*&lt;/strong&gt; mellom de 3 store cloud-tilbyderene i tillegg til å drifte selv om ønskelig.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;*&lt;/strong&gt; Kubernetes utvikles raskt, og funksjonalitet blir ofte ikke tilgjengelig på de ulike platformene samtidig.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;OppretteAKS&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;opprette-azure-kubernetes-cluster&quot;&gt;Opprette Azure Kubernetes Cluster&lt;/h2&gt;
&lt;p&gt;&lt;a id=&quot;Forbehold&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;forbehold&quot;&gt;Forbehold&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;Denne gjennomgangen tar utgangspunkt i dokumentasjonen på &lt;a href=&quot;https://docs.microsoft.com/en-us/azure/aks/kubernetes-walkthrough&quot;&gt;Microsoft.com&lt;/a&gt;. Å sette opp et Azure Kubernetes cluster fungerte ikke i starten av desember, men per dags dato, 23. desember, ser det ut til å fungere relativt bra. Men, oppgradering av cluster fra Kubernetes 1.7 til 1.8 fungerer for eksempel IKKE.&lt;/p&gt;
&lt;p&gt;AKS er i Preview og Azure jobber kontinuerlig med å gjøre AKS stabilt og støtte så mange Kubernetes-funksjoner som mulig. Amazon Web Services har tilsvarende en lukket invite-only Preview per dags dato mens de også jobber med stabilitet og funksjonalitet.&lt;/p&gt;
&lt;p&gt;Både Azure og AWS uttrykker forventning om at deres Kubernetes tjenester skal være klare for produksjonsmiljø ila 2018.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;Forberedelser&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;forberedelser&quot;&gt;Forberedelser&lt;/h3&gt;
&lt;p&gt;Du behøver Azure-CLI (versjon 2.0.21 eller nyere) for å utføre kommandoene:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://aka.ms/InstallAzureCliWindows&quot;&gt;Last ned Azure-CLI her&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.microsoft.com/en-us/cli/azure/install-azure-cli?view=azure-cli-latest&quot;&gt;Informasjon om Azure-CLI på MacOS og Linux finner du her&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Alle kommandoer gjøres i Windows PowerShell.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Azureinnlogging&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;azure-innlogging&quot;&gt;Azure innlogging&lt;/h3&gt;
&lt;p&gt;Logg på Azure:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az login&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Du får en link som du åpner i din browser samt en autentiseringskode. Skriv koden på nettsiden og &lt;code&gt;az login&lt;/code&gt; lagrer påloggingsinformasjonen slik at du ikke behøver å autentisere igjen på samme maskin.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;PS&lt;/strong&gt; Pålogingsinformasjonen lagres i &lt;code&gt;C:\Users\Brukernavn\.azure\&lt;/code&gt;. Du må selv passe på at ingen kopierer disse filene. Da får de full tilgang til din Azure konto.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;AktiverContainerService&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;aktiver-containerservice&quot;&gt;Aktiver ContainerService&lt;/h3&gt;
&lt;p&gt;Siden AKS er i Preview/Beta må du eksplisitt aktivere det for å få tilgang til &lt;code&gt;aks&lt;/code&gt; kommandoene.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az provider register -n Microsoft.ContainerService&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az provider show -n Microsoft.ContainerService&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;OpprettResourceGroup&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;opprett-en-resource-group&quot;&gt;Opprett en resource group&lt;/h3&gt;
&lt;p&gt;Her oppretter vi en resource group med navn “min_aks_rg” i Azure region West Europe.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az group create --name min_aks_rg --location westeurope&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Protip&lt;/strong&gt;
For å se en liste over tilgjengelige Azure regioner, bruk kommandoen &lt;code&gt;az account list-locations --output table&lt;/code&gt;. &lt;strong&gt;PS&lt;/strong&gt; Det kan hende AKS ikke er tilgjengelig i alle regioner enda.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;OpprettK8sCluster&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;opprette-kubernetes-cluster&quot;&gt;Opprette Kubernetes cluster&lt;/h3&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az aks create --resource-group min_aks_rg --name mitt_cluster --node-count 3 --generate-ssh-keys --node-vm-size Standard_B2s --node-osdisk-size 256 --kubernetes-version 1.8.2&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;--node-count&lt;/code&gt;
&lt;ul&gt;
&lt;li&gt;Antall vertsmaskiner tilgjengelig for å kjøre containers&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--generate-ssh-keys&lt;/code&gt;
&lt;ul&gt;
&lt;li&gt;Oppretter og outputter en SSH key som kan brukes for å SSHe direkte til vertsmaskinene.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--node-vm-size&lt;/code&gt;
&lt;ul&gt;
&lt;li&gt;Hvilken type Azure VM clusteret skal bestå av. For å se tilgjengelige størrelser bruk &lt;code&gt;az vm list-sizes -l westeurope --output table&lt;/code&gt; og &lt;a href=&quot;https://docs.microsoft.com/en-us/azure/virtual-machines/linux/sizes&quot;&gt;Microsofts nettsider.&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--node-osdisk-size&lt;/code&gt;
&lt;ul&gt;
&lt;li&gt;Disk størrelse på vertsmaskiner i GB. &lt;strong&gt;PS&lt;/strong&gt; Conteinere kan bli stoppet og flyttet til en annen host ved behov eller hvis en vertsmaskin forsvinner. Alle data lagret lokalt i conteineren blir da borte. Hvis en skal lagre ting permanent må en bruke PersistentVolumes og ikke lokal disk på vertsmaskin.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--kubernetes-version&lt;/code&gt;
&lt;ul&gt;
&lt;li&gt;Hvilken Kubernetes versjon som skal installeres. Azure installerer IKKE den siste versjonen som standard, og per dags dato fungerer ikke &lt;code&gt;az aks upgrade&lt;/code&gt; tilstrekkelig. Siste tilgjengelige versjon per dags dato er 1.8.2. Det er en fordel å bruke siste versjon da det skjer store forbedringer i Kubernetes fra versjon til versjon. Dokumentasjon er også mye bedre for nyere versjoner.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Lagre teksten som kommandoen spytter ut i en fil på en trygg plass. Den inneholder nøkler som kan brukes for å kople til clusteret med SSH. Selv om det i teorien ikke skal være nødvendig.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;InstallerKubectl&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;installer-kubectl&quot;&gt;Installer kubectl&lt;/h3&gt;
&lt;p&gt;&lt;code&gt;kubectl&lt;/code&gt; er klienten som gjør alle operasjoner mot ditt Kubernetes cluster. Azure CLI kan installere &lt;code&gt;kubectl&lt;/code&gt; for deg:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az aks install-cli&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Etter &lt;code&gt;kubectl&lt;/code&gt; er installert behøver vi å få påloggingsinformasjon slik at &lt;code&gt;kubectl&lt;/code&gt; kan kommunisere med Kubernetes clusteret.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az aks get-credentials --resource-group min_aks_rg --name mitt_cluster&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Påloggingsinformasjonen lagres i &lt;code&gt;C:\Users\Brukernavn\.kube\config&lt;/code&gt;. Hold disse filene hemmelig også.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Protip&lt;/strong&gt; Når en har flere ulike Kubernetes clusters kan en bytte hvilken &lt;code&gt;kubectl&lt;/code&gt; skal snakke til med &lt;code&gt;kubectl config get-contexts&lt;/code&gt; og &lt;code&gt;kubectl config set-context mitt_cluster&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;InspiserCluster&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;inspiser-cluster&quot;&gt;Inspiser cluster&lt;/h3&gt;
&lt;p&gt;For å se at clusteret og &lt;code&gt;kubectl&lt;/code&gt; virker begynner vi med noen kommandoer.&lt;/p&gt;
&lt;p&gt;Se alle vertsmaskiner og status:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;gt; kubectl get nodes&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAME                       STATUS    AGE       VERSION&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;aks-nodepool1-16970026-0   Ready     15m       v1.8.2&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;aks-nodepool1-16970026-1   Ready     15m       v1.8.2&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;aks-nodepool1-16970026-2   Ready     15m       v1.8.2&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Se alle tjenester, pods, deployments:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;gt; kubectl get all --all-namespaces&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAMESPACE     NAME                                          READY     STATUS    RESTARTS   AGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kube-system   po/kubernetes-dashboard-6fc8cf9586-frpkn      1/1       Running   0          3d&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAMESPACE     NAME                          CLUSTER-IP     EXTERNAL-IP     PORT(S)           AGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kube-system   svc/kubernetes-dashboard      10.0.161.132   &amp;lt;none&amp;gt;          80/TCP            3d&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAMESPACE     NAME                             DESIRED   CURRENT   UP-TO-DATE   AVAILABLE   AGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kube-system   deploy/kubernetes-dashboard      1         1         1            1           3d&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAMESPACE     NAME                                    DESIRED   CURRENT   READY     AGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kube-system   rs/kubernetes-dashboard-6fc8cf9586      1         1         1         3d&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Jeg har bare tatt et lite utdrag fra denne kommandoen. Du behøver ikke å forstå hva alle ressursene i &lt;code&gt;kube-system&lt;/code&gt; namespacet gjør. Det er hensikten at du skal slippe det når Microsoft står for management av selve clusteret.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Namespaces&lt;/strong&gt;
I Kubernetes er det noe som heter Namespaces. Ressurser i ett namespace har ikke automatisk tilgang til ressurser i et annet namespace. Tjenestene som Kubernetes selv benytter installeres i namespacet &lt;code&gt;kube-system&lt;/code&gt;. Kommandoen &lt;code&gt;kubectl&lt;/code&gt; viser deg vanligvis bare ressurser i &lt;code&gt;default&lt;/code&gt; namespace med mindre du spesifiserer &lt;code&gt;--all-namespaces&lt;/code&gt; eller &lt;code&gt;--namespace=xx&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;StarteNginx&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;starte-noen-nginx-containere&quot;&gt;Starte noen nginx containere&lt;/h3&gt;
&lt;blockquote&gt;
&lt;p&gt;En instans av en kjørende container kalles i Kubernetes for en &lt;strong&gt;Pod&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;nginx&lt;/code&gt; er en rask og fleksibel webserver.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Nå som clusteret er oppe å kjøre kan vi begynne å rulle ut tjenster og deployments på det.&lt;/p&gt;
&lt;p&gt;Vi begynner med å lage en Deployment bestående av 3 containere som alle kjører &lt;code&gt;nginx:mainline-alpine&lt;/code&gt; imaget fra &lt;a href=&quot;https://hub.docker.com/r/_/nginx/&quot;&gt;Docker hub&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;nginx-dep.yaml&lt;/strong&gt; ser slik ut:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apiVersion: apps/v1beta2&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kind: Deployment&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;metadata:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  name: nginx-deployment&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  labels:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    app: nginx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;spec:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  replicas: 3&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  selector:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    matchLabels:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      app: nginx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  template:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    metadata:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      labels:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        app: nginx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    spec:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      containers:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;      - name: nginx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        image: nginx:mainline-alpine&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        ports:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;        - containerPort: 80&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Last denne inn på clusteret med &lt;code&gt;kubectl create&lt;/code&gt;:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl create -f /attachments/2017-12-23-managed-kubernetes-on-azure/nginx-dep.yaml&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Denne kommandoen oppretter ressursene beskrevet i filen. &lt;code&gt;kubectl&lt;/code&gt; kan lese filer enten lokalt fra din maskin eller fra en URL.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Etter du har gjort endringer i en ressurs-definisjon (&lt;code&gt;.yaml&lt;/code&gt; fil) kan du oppdatere ressursene i clusteret med &lt;code&gt;kubectl replace -f ressurs.yaml&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Vi kan verifisere at Deployment er klar:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;gt; kubectl get deploy&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAME               DESIRED   CURRENT   UP-TO-DATE   AVAILABLE   AGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;nginx-deployment   3         3         3            3           10m&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Vi kan også hente de faktiske Pods som er startet:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;gt; kubectl get pods&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAME                                READY     STATUS    RESTARTS   AGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;nginx-deployment-569477d6d8-dqwx5   1/1       Running   0          10m&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;nginx-deployment-569477d6d8-xwzpw   1/1       Running   0          10m&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;nginx-deployment-569477d6d8-z5tfk   1/1       Running   0          10m&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Logger&lt;/strong&gt; Vi kan se logger fra én pod med &lt;code&gt;kubectl logs nginx-deployment-569477d6d8-xwzpw&lt;/code&gt;. Men siden vi i dette tilfellet ikke vet hvilken Pod som ender opp med å få innkommende forespørsler kan vi se logger fra alle Pods som har &lt;code&gt;app=nginx&lt;/code&gt; label: &lt;code&gt;kubectl logs -lapp=nginx&lt;/code&gt;. At vi her bruker &lt;code&gt;app=nginx&lt;/code&gt; har vi selv bestemt i &lt;code&gt;nginx-dep.yaml&lt;/code&gt; når vi satt &lt;code&gt;spec.template.metadata.labels: app: nginx&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;NginxService&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;gjøre-nginx-tilgjengelig-med-en-tjeneste&quot;&gt;Gjøre nginx tilgjengelig med en tjeneste&lt;/h3&gt;
&lt;p&gt;For å kommunisere med våre nye Pods behøver vi å opprette en tjeneste (&lt;strong&gt;Service&lt;/strong&gt;). En tjeneste består av en eller flere Pods som velges basert på ulike kriterier, blant annet hvilke labels de har og om Podene det gjelder er Running og Ready.&lt;/p&gt;
&lt;p&gt;Nå lager vi en tjeneste som ruter trafikk til alle Pods som har label &lt;code&gt;app: nginx&lt;/code&gt; og som lytter på port 80. I tillegg gjør vi tjenesten tilgjengelig via en LoadBalancer:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;nginx-svc.yaml&lt;/strong&gt; ser slik ut:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apiVersion: v1&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kind: Service&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;metadata:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  name: nginx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  labels:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    app: nginx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;spec:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  type: LoadBalancer&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  ports:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  - port: 80&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    name: http&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    targetPort: 80&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  selector:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    app: nginx&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Vi ber Kubernetes om å opprette tjeneten vår med &lt;code&gt;kubectl create&lt;/code&gt; som vanlig:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl create -f /attachments/2017-12-23-managed-kubernetes-on-azure/nginx-svc.yaml&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Deretter kan vi se hvilken IP-adresse tjenesten vår har fått av Azure:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;gt; kubectl get svc -w&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;NAME         CLUSTER-IP   EXTERNAL-IP     PORT(S)        AGE&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;nginx        10.0.24.11   13.95.173.255   80:31522/TCP   15m&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;PS&lt;/strong&gt; Det kan ta et par minutter for Azure å tildele tjenesten vår en Public IP, i mellomtiden vil det stå &lt;code&gt;&amp;lt;pending&amp;gt;&lt;/code&gt; under EXTERNAL-IP.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;En enkel &lt;strong&gt;Welcome to nginx&lt;/strong&gt; webside skal nå være tilgjengelig på &lt;a href=&quot;http://13.95.173.255&quot;&gt;http://13.95.173.255&lt;/a&gt; (&lt;em&gt;husk å bytt ut med din egen External-IP&lt;/em&gt;).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vi har nå en lastbalansert &lt;code&gt;nginx&lt;/code&gt; tjeneste med 3 servere klar til å ta imot trafikk.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For ordens skyld kan vi slette tjeneste og deployment etterpå:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl delete svc nginx&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;kubectl delete deploy nginx-deployment&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;ScaleCluster&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;skalere-cluster&quot;&gt;Skalere cluster&lt;/h3&gt;
&lt;p&gt;Hvis en ønsker å endre antall vertsmaskiner/noder som kjører Pods kan en gjøre det via Azure-CLI:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az aks scale --name mitt_cluster --resource-group min_aks_rg --node-count 5&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;For øyeblikket blir alle noder opprettet med samme størrelse som når clusteret ble opprettet. AKS vil antageligvis få støtte for &lt;a href=&quot;https://cloud.google.com/kubernetes-engine/docs/concepts/node-pools&quot;&gt;&lt;strong&gt;node-pools&lt;/strong&gt;&lt;/a&gt; i løpet av neste år. Da kan en opprette grupper av noder med forskjellig størrelse og operativsystem, både Linux og Windows.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;DeleteCluster&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;slette-cluster&quot;&gt;Slette cluster&lt;/h3&gt;
&lt;p&gt;En kan slette hele clusteret slik:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;az aks delete --name mitt_cluster --resource-group min_aks_rg --yes&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;Bonusmateriale&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;bonusmateriale&quot;&gt;Bonusmateriale&lt;/h2&gt;
&lt;p&gt;Her er litt bonusmateriale dersom du ønsker å gå enda litt videre med Kubernetes.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;HelmIntro&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;rulle-ut-tjenester-med-helm&quot;&gt;Rulle ut tjenester med Helm&lt;/h3&gt;
&lt;p&gt;&lt;a href=&quot;https://helm.sh/&quot;&gt;Helm&lt;/a&gt; er en pakke-behandler og et bibliotek av programvare som er klart for å rulles ut i et Kubernetes-cluster.&lt;/p&gt;
&lt;p&gt;Start med å laste ned &lt;a href=&quot;https://github.com/kubernetes/helm/releases&quot;&gt;Helm-klienten&lt;/a&gt;. Den henter påloggingsinformasjon osv fra samme sted som &lt;code&gt;kubectl&lt;/code&gt; automatisk.&lt;/p&gt;
&lt;p&gt;Installer Helm-serveren (Tiller) på Kubernetes clusteret og oppdater pakke-biblioteket:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;helm init&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;helm repo update&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Se tilgjengelige pakker (&lt;strong&gt;Charts&lt;/strong&gt;) med: &lt;code&gt;helm search&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;HelmMinecraft&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;rulle-ut-minecraft-med-helm&quot;&gt;Rulle ut MineCraft med Helm&lt;/h4&gt;
&lt;p&gt;La oss rulle ut en MineCraft installasjon på clusteret vårt, fordi vi kan :-)&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;helm install --name stian-sin --set minecraftServer.eula=true stable/minecraft&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;--set&lt;/code&gt; overstyrer en eller flere av standardverdiene som er satt i pakken. MineCraft pakken er laget slik at den ikke starter uten å ha sagt seg enig i brukervilkårene i variabelen &lt;code&gt;minecraftServer.eula&lt;/code&gt;. Alle variablene som kan overstyres i MineCraft pakken er &lt;a href=&quot;https://github.com/kubernetes/charts/blob/master/stable/minecraft/values.yaml&quot;&gt;dokumentert her&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Så venter vi litt på at Azure skal tildele en Public IP:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;gt; kubectl get svc -w&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;stian-sin-minecraft   10.0.237.0   13.95.172.192   25565:30356/TCP   3m&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Og vipps kan vi kople til Minecraft på &lt;code&gt;13.95.172.192:25565&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Kubernetes in MineCraft on Kubernetes&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./minecraft-k8s.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;KubernetesDashboard&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;kubernetes-dashboard&quot;&gt;Kubernetes Dashboard&lt;/h3&gt;
&lt;p&gt;Kubernetes har også et grafisk web-grensesnitt som gjør det litt lettere å se hvilke ressurser som er i clusteret, se logger og åpne remote-shell inne i en kjørende Pod, blant annet.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;gt; kubectl proxy&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;Starting to serve on 127.0.0.1:8001&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;kubectl&lt;/code&gt; krypterer og tunnelerer trafikken inn til Kubernetes’ API servere. Dashboardet er tilgjengelig på &lt;a href=&quot;http://127.0.0.1:8001/ui/&quot;&gt;http://127.0.0.1:8001/ui/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Kubernetes Dashboard&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./k8s-dash.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Konklusjon&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;konklusjon&quot;&gt;Konklusjon&lt;/h2&gt;
&lt;p&gt;Jeg håper du har fått mersmak for Kubernetes. Lærekurven kan være litt bratt i begynnelsen men det tar ikke så veldig lang tid før du er produktiv.&lt;/p&gt;
&lt;p&gt;Se på de &lt;a href=&quot;https://v1-8.docs.kubernetes.io/docs/tutorials/&quot;&gt;offisielle guidene på Kubernetes.io&lt;/a&gt; for å lære mer om hvordan du definerer forskjellige typer ressurser og tjenester for å kjøre på Kubernetes. &lt;strong&gt;PS: Det gjøres store endringer fra versjon til versjon så sørg for å bruke dokumentasjonen for riktig versjon!&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Kubernetes har også et veldig aktivt Slack-miljø på &lt;a href=&quot;http://slack.k8s.io/&quot;&gt;kubernetes.slack.com&lt;/a&gt;. Der er det også en kanal for norske Kubernetes brukere; &lt;strong&gt;#norw-users&lt;/strong&gt;.&lt;/p&gt;
</content:encoded><category>tech</category><category>kubernetes</category><category>azure</category></item><item><title>Next generation monitoring with OpenTSDB</title><link>https://blog.stian.omg.lol/p/next-generation-monitoring-with-opentsdb/</link><guid isPermaLink="true">https://blog.stian.omg.lol/p/next-generation-monitoring-with-opentsdb/</guid><description>Step by step guide on how to install a single-instance of OpenTSDB using the latest versions of Hadoop and HBase.</description><pubDate>Mon, 02 Jun 2014 19:56:40 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://blog.stian.omg.lol/_astro/2014-06-02-next-generation-monitoring-using-opentsdb.KQJcg8z7_Z1l4BJc.jpg&quot; width=&quot;1600&quot; height=&quot;914&quot; alt=&quot;&quot;&gt;&lt;/p&gt;
&lt;p&gt;Step by step guide on how to install a single-instance of OpenTSDB using the latest versions of Hadoop and HBase.&lt;/p&gt;
&lt;h3 id=&quot;2021-update-the-specific-tools-discussed-in-this-blog-post-should-be-considered-obsolete-by-todays-standards-you-should-investigate-prometheus-influxdb-and-timescaledb-for-your-monitoring-needs&quot;&gt;2021 Update: The specific tools discussed in this blog post should be considered obsolete by todays standards. You should investigate Prometheus, InfluxDB and TimescaleDB for your monitoring needs.&lt;/h3&gt;
&lt;p&gt;In this paper we will provide a step by step guide on how to install a single-instance of &lt;strong&gt;OpenTSDB&lt;/strong&gt; using the latest versions of the underlying technology, &lt;strong&gt;Hadoop&lt;/strong&gt; and &lt;strong&gt;HBase&lt;/strong&gt;. We will also provide some background on the state of existing monitoring solutions.
&lt;a id=&quot;Abstract&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Table of contents&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#Abstract&quot;&gt;Abstract&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Background&quot;&gt;Background&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#Performanceproblems&quot;&gt;Performance problems - Welcome to I/O-hell&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Scaling&quot;&gt;Scaling problems&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Loss&quot;&gt;Loss of detail&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#flexibility&quot;&gt;Lack of flexibility&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#revolution&quot;&gt;The monitoring revolution&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Debian&quot;&gt;Setting up a single node OpenTSDB instance on Debian 7 Wheezy&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#Hardware&quot;&gt;Hardware requirements&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Operating&quot;&gt;Operating system requirements&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#preparations&quot;&gt;Pre-setup preparations&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#java&quot;&gt;Installing java from packages&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#HBase&quot;&gt;Installing HBase&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#snappy&quot;&gt;Install snappy&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#native&quot;&gt;Building native libhadoop and libsnappy&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#ConfiguringHBase&quot;&gt;Configuring HBase&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#compression&quot;&gt;Testing HBase and compression&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#StartingHBase&quot;&gt;Starting HBase&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#InstallingOpenTSDB&quot;&gt;Installing OpenTSDB&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#ConfiguringOpenTSDB&quot;&gt;Configuring OpenTSDB&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#HBasetables&quot;&gt;Creating HBase tables&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#StartingOpenTSDB&quot;&gt;Starting OpenTSDB&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Feeding&quot;&gt;Feeding data into OpenTSDB&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#tcollector&quot;&gt;tcollector&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#peritus-tc-tools&quot;&gt;peritus-tc-tools&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#collectd-opentsdb&quot;&gt;collectd-opentsdb&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#MonitoringOpenTSDB&quot;&gt;Monitoring OpenTSDB&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Performancecomparison&quot;&gt;Performance comparison&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;#Collection&quot;&gt;Collection&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Storage&quot;&gt;Storage&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#Conclusion&quot;&gt;Conclusion&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a id=&quot;Background&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;background&quot;&gt;Background&lt;/h2&gt;
&lt;p&gt;Since its inception in 1999 &lt;a href=&quot;http://oss.oetiker.ch/rrdtool/&quot;&gt;&lt;strong&gt;rrdtool&lt;/strong&gt;&lt;/a&gt; (the underlying storage mechanism of once universal &lt;strong&gt;MRTG&lt;/strong&gt;) has been the base of many popular monitoring solutions; &lt;strong&gt;Cacti&lt;/strong&gt;, &lt;strong&gt;collectd&lt;/strong&gt;, &lt;strong&gt;Ganglia&lt;/strong&gt;, &lt;strong&gt;Munin&lt;/strong&gt;, &lt;strong&gt;Observium&lt;/strong&gt;, &lt;strong&gt;OpenNMS&lt;/strong&gt; and &lt;strong&gt;Zenoss&lt;/strong&gt;, to name a few.&lt;/p&gt;
&lt;p&gt;There are a number of problems with the current approach and we will highlight some of these here.&lt;/p&gt;
&lt;p&gt;Please note that this includes &lt;strong&gt;Graphite&lt;/strong&gt; and its backend &lt;strong&gt;Whisper&lt;/strong&gt;, which is based on the &lt;a href=&quot;http://graphite.readthedocs.org/en/0.9.10/whisper.html&quot;&gt;same basic design as rrdtool&lt;/a&gt; and has &lt;a href=&quot;http://dieter.plaetinck.be/on-graphite-whisper-and-influxdb.html&quot;&gt;some of the same limitations&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Performanceproblems&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;performance-problems---welcome-to-io-hell&quot;&gt;Performance problems - Welcome to I/O-hell&lt;/h3&gt;
&lt;p&gt;When MRTG and rrdtool was created the preservation of disk space was more important than preservation of disk operations and the default collection interval was 5 minutes (which many are still using). The way rrdtool is designed it requires quite a few random reads and writes per datapoint. It also re-reads, computes the average, and writes old data again according to the RRA rules defined which causes additional I/O load. In 2014 memory is cheap, disk storage is cheap and CPU is fairly cheap. Disk I/O operations (IOPS) however are still very expensive in terms of hardware. The recent maturing of SSD provides extreme amounts of IOPS for a reasonable price, but the drive sizes are fractional. The result is that in order to scale IOPS-wise you currently need many low-space SSDs to get the required space, or many low-IOPS spindle drives to get the required IOPS:&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;http://www.newegg.com/Product/Product.aspx?Item=N82E16820147251&quot;&gt;Samsung EVO 840 1TB SSD&lt;/a&gt; - 98.000 IOPS - 470 USD&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;http://www.newegg.com/Product/Product.aspx?Item=N82E16822148844&quot;&gt;Seagate Barracuda 3TB&lt;/a&gt; - 240 IOPS - 110 USD&lt;/p&gt;
&lt;p&gt;You would need $44.880 (408 drives) worth of spindle drives in order to match a single SSD drive in terms of I/O-performance. On the other hand a $2.000 array of spindle drives would get you a net ~54 TB of space. The cost of SSD to reach the same volume would be $25.380. Not to mention the cost of servers, power, provisioning, etc.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Note: This is the cheapest available bulk consumer drives and comparable OEM drives (&lt;a href=&quot;http://h30094.www3.hp.com/product/sku/10350615/mfg_partno/632494-B21&quot;&gt;SSD&lt;/a&gt;, &lt;a href=&quot;http://h30094.www3.hp.com/product/sku/10389145/mfg_partno/628061-B21&quot;&gt;spindle&lt;/a&gt;) for a HP server will be 6 to 30 times more expensive.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In rrdtool version 1.4, released in 2009, &lt;strong&gt;rrdcached&lt;/strong&gt; was introduced as a caching daemon for buffering multiple data updates and reducing the number of random I/O operations by writing several related datapoints in sequence. It took a couple of years before this new feature was implemented in most of the common open source monitoring solutions.&lt;/p&gt;
&lt;p&gt;For a good introduction into the internals of rrdtool/rrdcached updates and the problems with I/O scaling look at presentation by Sebastian Harl, &lt;a href=&quot;http://www.netways.de/index.php?id=2815&quot;&gt;How to Escape the I/O Hell&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Scaling&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;scaling-problems&quot;&gt;Scaling problems&lt;/h3&gt;
&lt;p&gt;Most of today’s monitoring systems do not easily scale-out. Scale-out, or scaling horizontally, is when you can add new nodes in response to increased load. Scaling up by replacing existing hardware with state-of-the-art hardware is both expensive and only buys you limited time before the next even more expensive necessary hardware upgrade. Many systems offer distributed polling but none offer the option of spreading out the disk load. For example; you can &lt;a href=&quot;http://community.zenoss.org/docs/DOC-2485&quot;&gt;scale Zenozz for High Availability&lt;/a&gt; but not performance.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Loss&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;loss-of-detail&quot;&gt;Loss of detail&lt;/h3&gt;
&lt;p&gt;Current RRD based systems will aggregate old data into averages in order to save storage space. Most technicians do not have the in depth knowledge in order to tune the rules for aggregation and will leave the default values as is. Using cacti as an example and looking at the &lt;a href=&quot;http://docs.cacti.net/manual:088:8_rrdtool#rrd_files&quot;&gt;cacti documentation&lt;/a&gt; we see that in a very short time, 2 months, data is averaged to a single data point PER DAY. For systems such as Internet backbones where traffic vary a lot from bottom (30% utilization for example) to peak (90% utilization for example) during a day only the average of 60% is shown in the graphs. This in turn makes troubleshooting by comparing old data difficult. It makes trending based on peaks/bottoms impossible and it may also lead to wrong or delayed strategic decisions on where to invest in added capacity.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;flexibility&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;lack-of-flexibility&quot;&gt;Lack of flexibility&lt;/h3&gt;
&lt;p&gt;In order to collect, store and graph new kinds of metrics an operator would need a certain level of programming skills and experience with the internals of the monitoring system. Adding new metrics to the systems would range from hours to weeks depending on the skill and experience of the operator. Creating new graphs based on existing metrics is also very difficult on most systems. And not within reach for the average operator.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;revolution&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;the-monitoring-revolution&quot;&gt;The monitoring revolution&lt;/h2&gt;
&lt;p&gt;We are currently at the beginning of a monitoring revolution. The advent of cloud computing and big data has created a need for measuring lots of metrics for thousands of machines at small intervals. This has sparked the creation of completely new monitoring components. One of the components where we now have improved alternatives is for efficient metric storage.&lt;/p&gt;
&lt;p&gt;The first is &lt;strong&gt;&lt;a href=&quot;http://opentsdb.net/&quot;&gt;OpenTSDB&lt;/a&gt;&lt;/strong&gt;, a “Scalable, Distributed, Time Series Database” that begun development at &lt;a href=&quot;https://www.stumbleupon.com/&quot;&gt;StumbleUpon&lt;/a&gt; in 2011 and aimed at solving some of the problems with existing monitoring systems. OpenTSDB is built in top of Apache HBase which is a scalable and performant database that builds on top of Apache Hadoop. Hadoop is a series of tools for building large and scalable distributed systems. Back in 2010 Facebook already had &lt;a href=&quot;http://hadoopblog.blogspot.no/2010/05/facebook-has-worlds-largest-hadoop.html&quot;&gt;2000 machines in a Hadoop cluster&lt;/a&gt; with 21PB (that is 21.000.000 GB) of combined storage.&lt;/p&gt;
&lt;p&gt;The second is an interesting newcommer, &lt;a href=&quot;http://influxdb.com/&quot;&gt;&lt;strong&gt;InfluxDB&lt;/strong&gt;&lt;/a&gt;, that began development in 2013 and has the goal of offering scalability and performance without the requirements of HBase/Hadoop.&lt;/p&gt;
&lt;p&gt;In addition to advances in performance these alternatives also decouple storage of metrics and display of graphs and abstract the interaction in simple and well-defined APIs. This makes it easy for developers to create improved frontends rapidly and this has already resulted in several very attractive open-source frontends such as &lt;strong&gt;&lt;a href=&quot;https://github.com/Ticketmaster/Metrilyx-2.0&quot;&gt;Metrilyx&lt;/a&gt;&lt;/strong&gt; (OpenTSDB), &lt;strong&gt;&lt;a href=&quot;http://grafana.org/&quot;&gt;Grafana&lt;/a&gt;&lt;/strong&gt; (InfluxDB, Graphite, &lt;a href=&quot;https://github.com/grafana/grafana/pull/211&quot;&gt;soon OpenTSDB&lt;/a&gt;), &lt;strong&gt;&lt;a href=&quot;http://www.statuswolf.com/&quot;&gt;StatusWolf&lt;/a&gt;&lt;/strong&gt; (OpenTSDB), &lt;strong&gt;&lt;a href=&quot;https://github.com/hakobera/influga&quot;&gt;Influga&lt;/a&gt;&lt;/strong&gt; (InfluxDB).&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Debian&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2 id=&quot;setting-up-a-single-node-opentsdb-instance-on-debian-7-wheezy&quot;&gt;Setting up a single node OpenTSDB instance on Debian 7 Wheezy&lt;/h2&gt;
&lt;p&gt;In the rest of this paper we will set up a single node OpenTSDB instance. OpenTSDB builds on top of HBase and Hadoop and scales to very large setups easily. But it also delivers substantial performance on a single node which is deployed in &lt;strong&gt;less than an hour&lt;/strong&gt;. There are plenty of guides on installing a Hadoop cluster but here we will focus on the natural first step of getting a single node running using &lt;strong&gt;recent releases&lt;/strong&gt; of the relevant software:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;OpenTSDB 2.0.0 - Released 2014-05-05&lt;/li&gt;
&lt;li&gt;HBase 0.98.2 - Released 2014-05-01&lt;/li&gt;
&lt;li&gt;Hadoop 2.4.0 - Released 2014-04-07&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;If you later require to deploy a larger cluster consider using a framework such as &lt;a href=&quot;http://www.cloudera.com/content/cloudera/en/products-and-services/cdh.html&quot;&gt;&lt;strong&gt;Cloudera CDH&lt;/strong&gt;&lt;/a&gt; or &lt;a href=&quot;http://hortonworks.com/hdp/&quot;&gt;&lt;strong&gt;Hortonworks HDP&lt;/strong&gt;&lt;/a&gt; which are open-source platforms which package Apache Hadoop components and provides a fully tested environment and easy-to-use graphical frontends for configuration and management. It is &lt;a href=&quot;http://opentsdb.net/setup-hbase.html&quot;&gt;recommended to have at least 5 machines&lt;/a&gt; in a HBase cluster supporting OpenTSDB.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;blockquote&gt;
&lt;p&gt;This guide assumes you are somewhat familiar with using a Linux shell/command prompt.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;Hardware&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;hardware-requirements&quot;&gt;Hardware requirements&lt;/h4&gt;
&lt;ul&gt;
&lt;li&gt;CPU cores: Max (Limit to 50% of your available CPU resources)&lt;/li&gt;
&lt;li&gt;RAM: Min 16 GB&lt;/li&gt;
&lt;li&gt;Disk 1 - OS: 10 GB - Thin provisioned&lt;/li&gt;
&lt;li&gt;Disk 2 - Data: 100 GB - Thin provisioned&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a id=&quot;Operating&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;operating-system-requirements&quot;&gt;Operating system requirements&lt;/h4&gt;
&lt;p&gt;This guide is based on a recently installed Debian 7 Wheezy &lt;strong&gt;64bit&lt;/strong&gt; installed without any extra packages. See the &lt;a href=&quot;https://www.debian.org/releases/stable/amd64/&quot;&gt;official documentation&lt;/a&gt; for more information.&lt;/p&gt;
&lt;p&gt;All commands are entered as &lt;strong&gt;root&lt;/strong&gt; user unless otherwise noted.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;preparations&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;pre-setup-preparations&quot;&gt;Pre-setup preparations&lt;/h4&gt;
&lt;p&gt;We start by installing a few tools that we will need later.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apt-get install wget make gcc g++ cmake maven&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Create a new ext3 partition on the data disk &lt;strong&gt;/dev/sdb&lt;/strong&gt;:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;(echo &quot;n&quot;; echo &quot;p&quot;; echo &quot;&quot;; echo &quot;&quot;; echo &quot;&quot;; echo &quot;t&quot;; echo &quot;83&quot;; echo &quot;w&quot;) | fdisk /dev/sdb&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;mkfs.ext3 /dev/sdb1&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;ext3 is the &lt;a href=&quot;https://wiki.apache.org/hadoop/DiskSetup&quot;&gt;recommended filesystem for Hadoop&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Create a mountpoint &lt;strong&gt;/mnt/data1&lt;/strong&gt; and add it to the file system table and mount the disk:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;mkdir /mnt/data1&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;echo &quot;/dev/sdb1     /mnt/data1    ext3    auto,noexec,noatime,nodiratime   0   1&quot; | tee -a /etc/fstab&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;mount /mnt/data1&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;Using &lt;strong&gt;noexec&lt;/strong&gt; for the data partition will increase security as nothing on the data partition will be allowed to ever execute.
Using &lt;strong&gt;noatime&lt;/strong&gt; and &lt;strong&gt;nodiratime&lt;/strong&gt; increases performance since the read access timestamps are not updated on every file access.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;java&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;installing-java-from-packages&quot;&gt;Installing java from packages&lt;/h4&gt;
&lt;p&gt;Installing java on Linux can be quite challenging due to licensing issues, but thanks to the guys over at &lt;a href=&quot;https://launchpad.net/&quot;&gt;Launchpad.net&lt;/a&gt; who are providing a repository with a custom java package this can now be done quite easy.&lt;/p&gt;
&lt;p&gt;We start by adding the launchpad java repository to our &lt;em&gt;&lt;strong&gt;/etc/apt/sources.list&lt;/strong&gt;&lt;/em&gt; file:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;echo &quot;deb http://ppa.launchpad.net/webupd8team/java/ubuntu precise main&quot; | tee -a /etc/apt/sources.list&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;echo &quot;deb-src http://ppa.launchpad.net/webupd8team/java/ubuntu precise main&quot; | tee -a /etc/apt/sources.list&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Add the signing key and download information from the new repository:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apt-key adv --keyserver hkp://keyserver.ubuntu.com:80 --recv-keys EEA14886&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apt-get update&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Run the java installer:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apt-get install oracle-java7-installer&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Follow the instructions on screen to complete the Java 7 installation.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;HBase&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;installing-hbase&quot;&gt;Installing HBase&lt;/h3&gt;
&lt;p&gt;OpenTSDB has its own HBase installation tutorial &lt;a href=&quot;http://opentsdb.net/setup-hbase.html&quot;&gt;here&lt;/a&gt;. It is very brief and does not use the latest versions or snappy compression.&lt;/p&gt;
&lt;p&gt;Download and unpack HBase:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cd /opt&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;wget http://apache.vianett.no/hbase/hbase-0.98.2/hbase-0.98.2-hadoop2-bin.tar.gz&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;tar xvfz hbase-0.98.2-hadoop2-bin.tar.gz&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;export HBASEDIR=`pwd`/hbase-0.98.2-hadoop2/&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Increase the system-wide limitations of open files and processes from the default of 1000 to 32000 by adding a few lines to &lt;em&gt;&lt;strong&gt;/etc/security/limits.conf&lt;/strong&gt;&lt;/em&gt;:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;echo &quot;root    -               nofile  32768&quot; | tee -a /etc/security/limits.conf&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;echo &quot;root    soft/hard       nproc   32000&quot; | tee -a /etc/security/limits.conf&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;echo &quot;*    -               nofile  32768&quot; | tee -a /etc/security/limits.conf&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;echo &quot;*    soft/hard       nproc   32000&quot; | tee -a /etc/security/limits.conf&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The settings above will only take effect if we also add a line to &lt;em&gt;&lt;strong&gt;/etc/pam.d/common-session&lt;/strong&gt;&lt;/em&gt;:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;echo &quot;session required  pam_limits.so&quot; | tee -a /etc/pam.d/common-session&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;snappy&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;install-snappy&quot;&gt;Install snappy&lt;/h4&gt;
&lt;p&gt;&lt;a href=&quot;https://code.google.com/p/snappy/&quot;&gt;Snappy&lt;/a&gt; is a compression algorithm that values speed over compression ratio and this makes it a good choice for high throughput applications such as Hadoop/HBase. Due to licensing issues Snappy does not ship with HBase and need to be installed on top.&lt;/p&gt;
&lt;p&gt;The installation process is a bit complicated and has caused headache for many people (me included). Here we will show a method of installing snappy and getting it to work with the latest version of HBase and Hadoop.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Compression algorithms in HBase&lt;/strong&gt;
Compression is the method of reducing the size of a file or text without losing any of the contents. There are many compression algorithms available and some focus on being able to create the smallest compressed file at the cost of time and CPU usage while other achieve &lt;em&gt;reasonable&lt;/em&gt; compression ratio while being very fast.
Out of the box HBase supports gz(gzip/zlib), snappy and lzo. Only gz is included due to licensing issues.
Unfortunately gz is a slow and costly algorithm compared to snappy and lzo. In a test performed by Yahoo (see &lt;a href=&quot;http://www.slideshare.net/Hadoop_Summit/singh-kamat-june27425pmroom210c&quot;&gt;slides here&lt;/a&gt;, page 8) gz achieves 64% compression in 32 seconds. lzo 47% in 4.8 seconds and snappy 42% in 4.0 seconds. lz4 is another protocol &lt;a href=&quot;http://search-hadoop.com/m/KFLWV1PFVhp1&quot;&gt;considered for inclusion&lt;/a&gt; that is even faster (2.4 seconds) but requires much more memory.
&lt;em&gt;For more information look at the &lt;a href=&quot;https://hbase.apache.org/book/compression.html&quot;&gt;Apache HBase Handbook - Appendix C - Compression&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a id=&quot;native&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;building-native-libhadoop-and-libsnappy&quot;&gt;Building native libhadoop and libsnappy&lt;/h4&gt;
&lt;p&gt;In order to use compression we need the common Hadoop library, libhadoop.so, and the snappy library, libsnappy.so. HBase ships without libhadoop.so and the libhadoop.so that ships in the Hadoop Package is only for 32 bit OS. So we need to compile these files ourself.&lt;/p&gt;
&lt;p&gt;Start by downloading and installing ProtoBuf. Hadoop requres version 2.5+ which is not available as a Debian package unfortunately.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;wget --no-check-certificate https://protobuf.googlecode.com/files/protobuf-2.5.0.tar.gz&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;tar zxvf protobuf-2.5.0.tar.gz&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cd protobuf-2.5.0&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;./configure; make; make install&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;export LD_LIBRARY_PATH=/usr/local/lib/&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Download and compile Hadoop:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apt-get install zlib1g-dev&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;wget http://apache.uib.no/hadoop/common/hadoop-2.4.0/hadoop-2.4.0-src.tar.gz&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;tar zxvf hadoop-2.4.0-src.tar.gz&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cd hadoop-2.4.0-src/hadoop-common-project/&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;mvn package -Pdist,native -Dskiptests -Dtar -Drequire.snappy -DskipTests&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Copy the newly compiled native libhadoop library into /usr/local/lib, then create the folder in which HBase looks for it and create a shortcut from there to /usr/local/lib/libhadoop.so:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cp hadoop-common/target/native/target/usr/local/lib/libhadoop.* /usr/local/lib&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;mkdir -p $HBASEDIR/lib/native/Linux-amd64-64/&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cd $HBASEDIR/lib/native/Linux-amd64-64/&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;ln -s /usr/local/lib/libhadoop.so* .&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Install snappy from Debian packages:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apt-get install libsnappy-dev&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;ConfiguringHBase&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;configuring-hbase&quot;&gt;Configuring HBase&lt;/h4&gt;
&lt;p&gt;Now we need to do some basic configuration before we can start HBase. The configuration files are in $HBASEDIR/conf/.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;hbase-env.sh&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;confhbase-envsh&quot;&gt;&lt;strong&gt;conf/hbase-env.sh&lt;/strong&gt;&lt;/h4&gt;
&lt;p&gt;A shell script setting various environment variables related to how HBase and Java should behave. The file contains a lot of options and they are all documented by comments so feel free to look around in it.&lt;/p&gt;
&lt;p&gt;Start by setting the JAVA_HOME, which points to where Java is installed:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;export JAVA_HOME=/usr/lib/jvm/java-7-oracle/&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Then increase the size of the &lt;a href=&quot;http://pubs.vmware.com/vfabric52/index.jsp?topic=/com.vmware.vfabric.em4j.1.2/em4j/conf-heap-management.html&quot;&gt;Java Heap&lt;/a&gt; from the default of 1000 which is a bit low:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;export HBASE_HEAPSIZE=8000&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;hbase-site-xml&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;confhbase-sitexml&quot;&gt;&lt;strong&gt;conf/hbase-site.xml&lt;/strong&gt;&lt;/h4&gt;
&lt;p&gt;An XML file containing HBase specific configuration parameters.&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;lt;configuration&amp;gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;   &amp;lt;property&amp;gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &amp;lt;name&amp;gt;hbase.rootdir&amp;lt;/name&amp;gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &amp;lt;value&amp;gt;/mnt/data1/hbase&amp;lt;/value&amp;gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;lt;/property&amp;gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;lt;property&amp;gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &amp;lt;name&amp;gt;hbase.zookeeper.property.dataDir&amp;lt;/name&amp;gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;    &amp;lt;value&amp;gt;/mnt/data1/zookeeper&amp;lt;/value&amp;gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;  &amp;lt;/property&amp;gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;&amp;lt;/configuration&amp;gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;compression&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;testing-hbase-and-compression&quot;&gt;Testing HBase and compression&lt;/h4&gt;
&lt;p&gt;Now that we have installed snappy and configured HBase we can verify that HBase is working and that the compression is loaded by doing:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;$HBASEDIR/bin/hbase org.apache.hadoop.hbase.util.CompressionTest /tmp/test.txt snappy&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This should output some lines with information and end with &lt;strong&gt;SUCCESS&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;StartingHBase&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;starting-hbase&quot;&gt;Starting HBase&lt;/h4&gt;
&lt;p&gt;HBase ships with scripts for starting and stopping it, namely start-hbase.sh and stop-hbase.sh. You start HBase with&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;$HBASEDIR/bin/start-hbase.sh&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Then look at the log to ensure it has started without any serious errors:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;tail -fn100 $HBASEDIR/bin/../logs/hbase-root-master-opentsdb.log&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;If you want HBase to start automatically on boot you can use a process management tool such as &lt;a href=&quot;http://mmonit.com/monit/&quot;&gt;Monit&lt;/a&gt; or simply put it in &lt;em&gt;&lt;strong&gt;/etc/rc.local&lt;/strong&gt;&lt;/em&gt;:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;/opt/hbase-0.98.2-hadoop2/bin/start-hbase.sh&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;InstallingOpenTSDB&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;installing-opentsdb&quot;&gt;Installing OpenTSDB&lt;/h3&gt;
&lt;p&gt;Start by installing gnuplot, which is used by the native webui to draw graphs:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;apt-get install gnuplot&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Then download and install OpenTSDB:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;wget https://github.com/OpenTSDB/opentsdb/releases/download/v2.0.0/opentsdb-2.0.0_all.deb&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;dpkg -i opentsdb-2.0.0_all.deb&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;ConfiguringOpenTSDB&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;configuring-opentsdb&quot;&gt;Configuring OpenTSDB&lt;/h4&gt;
&lt;p&gt;The configuration file is &lt;em&gt;&lt;strong&gt;/etc/opentsdb/opentsdb.conf&lt;/strong&gt;&lt;/em&gt;. It has some of the basic configuration parameters but not nearly all of them. &lt;a href=&quot;http://opentsdb.net/docs/build/html/user_guide/configuration.html&quot;&gt;Here is the official documentation with all configuration parameters&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The defaults are reasonable but we need to make a few tweaks, the first is to add this:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;tsd.core.auto_create_metrics = true&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This will make OpenTSDB accept previously unseen metrics and add them to the database. This is very useful in the beginning when feeding data into OpenTSDB. Without this you will have to use the command &lt;em&gt;&lt;strong&gt;mkmetric&lt;/strong&gt;&lt;/em&gt; for each metric you will store and get errors that might be hard to trace if the metric you create do not match what is actually sent.&lt;/p&gt;
&lt;p&gt;Then we will add support for chunked requests via the HTTP API:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;tsd.http.request.enable_chunked = true&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;tsd.http.request.max_chunk = 16000&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Some tools and plugins (such as our own &lt;a href=&quot;https://github.com/PeritusConsulting/collectd-opentsdb&quot;&gt;improved collectd to OpenTSDB plugin&lt;/a&gt;) send multiple data points in a single HTTP request for increased efficiency and requires this setting to be enabled.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;HBasetables&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;creating-hbase-tables&quot;&gt;Creating HBase tables&lt;/h4&gt;
&lt;p&gt;Before we start OpenTSDB we need to create the necessary tables in HBase:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;env COMPRESSION=SNAPPY HBASE_HOME=$HBASEDIR /usr/share/opentsdb/tools/create_table.sh&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;a id=&quot;StartingOpenTSDB&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;starting-opentsdb&quot;&gt;Starting OpenTSDB&lt;/h4&gt;
&lt;p&gt;Since version 2.0.0 OpenTSDB ships as a Debian package and includes SysV init scripts. To start OpenTSDB as a daemon running in the background we run:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;service opentsdb start&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And then check the logs for any errors or other relevant information:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;tail -f /var/log/opentsdb/opentsdb.log&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;If the server is started successfully the last line of the log should say:&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;13:42:30.900 INFO  [TSDMain.main] - Ready to serve on /0.0.0.0:4242&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And you can now browse to your new OpenTSDB in a browser using &lt;a href=&quot;http://hostname:4242&quot;&gt;http://hostname:4242&lt;/a&gt; !&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Feeding&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;feeding-data-into-opentsdb&quot;&gt;Feeding data into OpenTSDB&lt;/h3&gt;
&lt;p&gt;It is not within the scope of this paper to go into details about how to feed data into OpenTSDB but we will give a quick introduction here to get you started.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A note on metric naming in OpenTSDB&lt;/strong&gt;
Each datapoint has a metric name such as &lt;em&gt;&lt;strong&gt;df.bytes.free&lt;/strong&gt;&lt;/em&gt; and one or more tags such as &lt;em&gt;&lt;strong&gt;host=server1&lt;/strong&gt;&lt;/em&gt; and &lt;em&gt;&lt;strong&gt;mount=/mnt/data1&lt;/strong&gt;&lt;/em&gt;. This is closer to the proposed &lt;a href=&quot;http://metrics20.org/&quot;&gt;Metrics 2.0&lt;/a&gt; standard for naming metrics than the traditional naming of &lt;em&gt;&lt;strong&gt;df.bytes.free.server1.mnt-data&lt;/strong&gt;&lt;/em&gt;. This makes it possible to create aggregates across tags and combine data easily using the tags.
OpenTSDB stores each datapoint with a given metric and tags in one HBase row per hour. But due to a HBase issue it still has to scan every row that matches the metric, ignoring the tags. Even though it will only return the data also matching the tags. This results in very much data being read and it will be very slow to read if there is a large number of data points for a given metric. The default for the collectd-opentsdb plugin is to use the read plugin name as metric, and other values as tags. In my case this results in 72.000.000 datapoints per hour for this metric. When generating a graph all of this data has to be read and evaluated before drawing a graph. 24 hours of data is over 1.7 billion datapoints for this single metric and results in a read performance of 5-15 &lt;strong&gt;minutes&lt;/strong&gt; for a simple graph.
A solution to this is to use &lt;em&gt;shift-to-metric&lt;/em&gt;, as &lt;a href=&quot;http://opentsdb.net/docs/build/html/user_guide/writing.html&quot;&gt;mentioned in the OpenTSDB user guide&lt;/a&gt;. Shift-to-metric is simply moving one or more data identifiers from tags to the metric in order to reduce the cardinality (number of values) for a metric, and hence the time required to read out the data we want. We have modified the collectd-opentsdb java plugin in order to shift the tags to metrics, and this increases read-performance by ~1000x down to 10-100ms. Read the section about collectd below for more information on our modified plugin.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id=&quot;tcollector&quot;&gt;tcollector&lt;/h4&gt;
&lt;p&gt;&lt;a href=&quot;http://opentsdb.net/docs/build/html/user_guide/utilities/tcollector.html&quot;&gt;tcollector&lt;/a&gt; is the default agent for collecting and sending data from a Linux server to a OpenTSDB server. It is based on Python and plugins / addons can be written in any language. It ships with the most common plugins to collect information about disk usage and performance, cpu and memory statistics and also for some specific systems such as mysql, mongodb, riak, varnish, postgresql and others. tcollector is very lightweight and features advanced de-duplication in order to reduce unneeded network traffic.&lt;/p&gt;
&lt;p&gt;The commands for installing dependencies and downloading tcollector are&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;aptitude install git python&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;cd /opt&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span&gt;git clone git://github.com/OpenTSDB/tcollector.git&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Configuration is in the startup script &lt;em&gt;&lt;strong&gt;tcollector/startstop&lt;/strong&gt;&lt;/em&gt;, you will need to uncomment and set the value of TSD_HOST to point to your OpenTSDB server.&lt;/p&gt;
&lt;p&gt;To start it run&lt;/p&gt;
&lt;pre class=&quot;astro-code astro-code-themes github-light github-dark&quot; style=&quot;--shiki-light:#24292e;--shiki-dark:#e1e4e8;--shiki-light-bg:#fff;--shiki-dark-bg:#24292e; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;plaintext&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span&gt;/opt/tcollector/startstop start&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This is also the command you want to add to &lt;em&gt;&lt;strong&gt;/etc/rc.local&lt;/strong&gt;&lt;/em&gt; in order to have the agent automatically start at boot. Logfiles are saved in &lt;em&gt;&lt;strong&gt;/var/log/tcollector.log&lt;/strong&gt;&lt;/em&gt; and they are rotated automatically.&lt;/p&gt;
&lt;h4 id=&quot;peritus-tc-tools&quot;&gt;peritus-tc-tools&lt;/h4&gt;
&lt;p&gt;We have developed a set of &lt;strong&gt;tcollector&lt;/strong&gt; plugins for collecting statistics from&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&quot;https://www.isc.org/downloads/dhcp/&quot;&gt;ISC DHCPd server&lt;/a&gt;&lt;/strong&gt;, about number of DHCP events and DHCP pool sizes&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&quot;http://www.opensips.org/&quot;&gt;OpenSIPS&lt;/a&gt;&lt;/strong&gt;, total number of subscribers and registered user agents&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&quot;http://atmail.com/&quot;&gt;Atmail&lt;/a&gt;&lt;/strong&gt;, number of users, admins, sent and received emails, logins and errors&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;As well as a high performance replacement for &lt;strong&gt;&lt;a href=&quot;http://oss.oetiker.ch/smokeping/&quot;&gt;smokeping&lt;/a&gt;&lt;/strong&gt; called &lt;strong&gt;tc-ping&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;These plugins are available for download from our &lt;strong&gt;&lt;a href=&quot;https://github.com/PeritusConsulting/peritus-tc-tools&quot;&gt;GitHub page&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;
&lt;h4 id=&quot;collectd-opentsdb&quot;&gt;collectd-opentsdb&lt;/h4&gt;
&lt;p&gt;&lt;a href=&quot;http://collectd.org/&quot;&gt;collectd&lt;/a&gt; is the &lt;em&gt;system statistics collection daemon&lt;/em&gt; and is a widely used system for collecting metrics from various sources. There are several options for sending data from collectd to OpenTSDB but one way that works well is to use the &lt;a href=&quot;https://github.com/auxesis/collectd-opentsdb&quot;&gt;collectd-opentsdb java write plugin&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Since collectd is a generic metric collection tool the original collectd-opentsdb plugin will use the plugin name (such as &lt;strong&gt;snmp&lt;/strong&gt;) as the metric, and use tags such as &lt;strong&gt;host=servername&lt;/strong&gt;, &lt;strong&gt;plugin_instance=ifHcInOctets&lt;/strong&gt; and &lt;strong&gt;type_instance=FastEthernet0/1&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;As mentioned in the &lt;em&gt;&lt;strong&gt;note on metric naming in OpenTSDB&lt;/strong&gt;&lt;/em&gt; this can be very inefficient when data needs to be read again resulting in read performance potentially thousands of times slower than optimal (&amp;lt;100ms). To alleviate this we have modified the original collectd-opentsdb plugin to store all metadata as part of the metric. This gives metric names such as ifHCInBroadcastPkts.sw01.GigabitEthernet0 and very good read performance.&lt;/p&gt;
&lt;p&gt;The modified collectd-opentsdb plugin can be downloaded from our &lt;a href=&quot;https://github.com/PeritusConsulting/collectd-opentsdb&quot;&gt;GitHub repository&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;MonitoringOpenTSDB&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;monitoring-opentsdb&quot;&gt;Monitoring OpenTSDB&lt;/h4&gt;
&lt;p&gt;To monitor OpenTSDB itself install tcollector as described above on the OpenTSDB server and set &lt;em&gt;&lt;strong&gt;TSD_HOST&lt;/strong&gt;&lt;/em&gt; to &lt;em&gt;&lt;strong&gt;localhost&lt;/strong&gt;&lt;/em&gt; in &lt;em&gt;&lt;strong&gt;/opt/tcollector/startstop&lt;/strong&gt;&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;You can then go to &lt;a href=&quot;http://opentsdb-server:4242/#start=1h-ago&amp;amp;end=1s-ago&amp;amp;m=sum:rate:tsd.rpc.received%7Btype=%5C*%7D&amp;amp;o=&amp;amp;yrange=%5B0:%5D&amp;amp;wxh=1200x600&quot;&gt;http://opentsdb-server:4242/#start=1h-ago&amp;amp;end=1s-ago&amp;amp;m=sum:rate:tsd.rpc.received%7Btype=\*%7D&amp;amp;o=&amp;amp;yrange=%5B0:%5D&amp;amp;wxh=1200x600&lt;/a&gt; to view a graph of amount of data received in the last hour.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Performancecomparison&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3 id=&quot;performance-comparison&quot;&gt;Performance comparison&lt;/h3&gt;
&lt;p&gt;Lastly we include a little performance comparison between the latest version of OpenTSDB+HBase+Hadoop, a previous version of OpenTSDB+HBase+Hadoop that we have used for a while as well as rrdcached which ran in production for 4 years at a client.&lt;/p&gt;
&lt;p&gt;The workload is gathering and storing metrics from 150 Cisco switches with 8200 ports/interfaces every 5 seconds. This equals about 15.000 points per second.&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Figure 1 - Data received by OpenTSDB per second&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./Figure1.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Collection&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;collection&quot;&gt;Collection&lt;/h4&gt;
&lt;p&gt;Even though it is not the primary focus, we include some data about collection performance for completeness. Collection is done using the latest version of &lt;a href=&quot;http://collectd.org/&quot;&gt;collectd&lt;/a&gt; and the builtin SNMP plugin.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;NB #1:&lt;/strong&gt; There is a &lt;a href=&quot;https://github.com/collectd/collectd/issues/610&quot;&gt;memory leak&lt;/a&gt; in the way collectd’s SNMP plugin uses the underlying libsnmp library and you might need to schedule a restart of the collectd service as a workaround for that if handling large workloads.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;NB #2:&lt;/strong&gt; Due to &lt;a href=&quot;http://comments.gmane.org/gmane.comp.monitoring.collectd/5061&quot;&gt;limitations in the libnetsnmp library&lt;/a&gt; you will run into problems if polling many (1000+) devices with a single collectd instance. A workaround is to run multiple collectd instances with fewer hosts. &lt;a href=&quot;https://github.com/collectd/collectd/issues/610&quot;&gt;memory leak&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Figure 2 shows that collection through SNMP polling consumes about 2200Mhz. We optimized some of the data types and definitions in collectd when moving to OpenTSDB and achieved a 20% performance increase in the polling as seen in Figure 3.&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Figure 2 - CPU Usage - SNMP polling and writing to RRDcached&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./Figure2.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Figure 3 - CPU Usage - SNMP polling and sending to OpenTSDB&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./Figure3.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;Writing to the native rrdcached write plugin consumes 1300Mhz while our modified collectd-opentsdb plugin consumes 1450Mhz. It is probably possible to create a much more efficient write plugin with more advanced knowledge of concurrency and using a lower level language such as C.&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Storage&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;storage&quot;&gt;Storage&lt;/h4&gt;
&lt;p&gt;When considering storage performance we will look at CPU usage and disk IOPS since these are the primary drivers of cost in today’s datacenters.&lt;/p&gt;
&lt;h4 id=&quot;collectd--rrdcached&quot;&gt;collectd + rrdcached&lt;/h4&gt;
&lt;p&gt;CPU usage - 1300Mhz, see Figure 2 above.&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Figure 4 - Disk write IOPS - Fluctuating between 10 and 170 IOPS during the 1 hour flush period.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./Figure4.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;h4 id=&quot;opentsdb--hbase-096--hadoop-1&quot;&gt;OpenTSDB + Hbase 0.96 + Hadoop 1&lt;/h4&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Figure 5 - CPU usage - 1700Mhz baseline with peaks of 7000Mhz during Java Garbage Collection (GC) (untuned).&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./Figure5.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Figure 6 - Disk write IOPS - 5 IOPS average with peaks of 25 IOPS during Java GC. We also see that disk read IOPS are much higher and this is due to regular compaction of the database and can be tuned. Reads in general can be reduced by increasing caching with more RAM if necessary.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./Figure6.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;h4 id=&quot;opentsdb--hbase-098--hadoop-2&quot;&gt;OpenTSDB + HBase 0.98 + Hadoop 2&lt;/h4&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Figure 7 - CPU usage - 1200Mhz baseline with peaks of 5000-6000Mhz during Java GC (untuned).&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./Figure7.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;quot;alt&amp;quot;:&amp;quot;Figure 8 - Disk write IOPS - &lt; 5 IOPS average with peaks of 25 IOPS during Java GC. Much less read IOPS during compaction compared to HBase 0.96.&amp;quot;,&amp;quot;src&amp;quot;:&amp;quot;./Figure8.png&amp;quot;,&amp;quot;index&amp;quot;:0}&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a id=&quot;Conclusion&quot;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h4&gt;
&lt;p&gt;Even without tuning, a single instance OpenTSDB installation is able to handle significant amounts of data before running into IO problems. This comes at a cost of CPU, currently OpenTSDB will consume &amp;gt; 300% the amount of CPU cycles compared to rrdcached for storage. But this is offset by a 85-95% reduction in disk load. In absolute terms for our particular set up (one 2 year old HP DL360p Gen8 running VMware vSphere 5.5) CPU usage increased from 15% to 25% while reducing IOPS load from 70% to &amp;lt; 10%.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Fine tuning of parameters (such as Java GC) as well as detailed analysis of memory usage is outside the scope of this brief paper and detailed information may be found elsewhere (&lt;a href=&quot;https://hbase.apache.org/book/performance.html&quot;&gt;51&lt;/a&gt;,&lt;a href=&quot;http://www.oracle.com/technetwork/java/javase/gc-tuning-6-140523.html&quot;&gt;52&lt;/a&gt;,&lt;a href=&quot;http://www.cubrid.org/blog/textyle/428187&quot;&gt;53&lt;/a&gt;) for those interested.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
</content:encoded><category>tech</category><category>observability</category></item></channel></rss>