Commit 6e9b0190 authored Mar 26, 2024 by Jakub Kicinski

net: remove gfp_mask from napi_alloc_skb()



__napi_alloc_skb() is napi_alloc_skb() with the added flexibility
of choosing gfp_mask. This is a NAPI function, so GFP_ATOMIC is
implied. The only practical choice the caller has is whether to
set __GFP_NOWARN. But that's a false choice, too, allocation failures
in atomic context will happen, and printing warnings in logs,
effectively for a packet drop, is both too much and very likely
non-actionable.

This leads me to a conclusion that most uses of napi_alloc_skb()
are simply misguided, and should use __GFP_NOWARN in the first
place. We also have a "standard" way of reporting allocation
failures via the queue stat API (qstats::rx-alloc-fail).

The direct motivation for this patch is that one of the drivers
used at Meta calls napi_alloc_skb() (so prior to this patch without
__GFP_NOWARN), and the resulting OOM warning is the top networking
warning in our fleet.

Reviewed-by: Alexander Lobakin <aleksander.lobakin@intel.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Link: https://lore.kernel.org/r/20240327040213.3153864-1-kuba@kernel.org


Signed-off-by: Jakub Kicinski <kuba@kernel.org>

parent 49d665b8

Documentation/mm/page_frags.rst

+1 −1

Original line number	Diff line number	Diff line
		@@ -25,7 +25,7 @@ to be disabled when executing the fragment allocation.
		The network stack uses two separate caches per CPU to handle fragment
		allocation. The netdev_alloc_cache is used by callers making use of the
		netdev_alloc_frag and __netdev_alloc_skb calls. The napi_alloc_cache is
		used by callers of the __napi_alloc_frag and __napi_alloc_skb calls. The
		used by callers of the __napi_alloc_frag and napi_alloc_skb calls. The
		main difference between these two calls is the context in which they may be
		called. The "netdev" prefixed functions are usable in any context as these
		functions will disable interrupts, while the "napi" prefixed functions are

Documentation/translations/zh_CN/mm/page_frags.rst

+1 −1

Original line number	Diff line number	Diff line
		@@ -25,7 +25,7 @@ sk_buff->head使用，或者用于skb_shared_info的 “frags” 部分。

		网络堆栈在每个CPU使用两个独立的缓存来处理碎片分配。netdev_alloc_cache被使用
		netdev_alloc_frag和__netdev_alloc_skb调用的调用者使用。napi_alloc_cache
		被调用__napi_alloc_frag和__napi_alloc_skb的调用者使用。这两个调用的主要区别是
		被调用__napi_alloc_frag和napi_alloc_skb的调用者使用。这两个调用的主要区别是
		它们可能被调用的环境。“netdev” 前缀的函数可以在任何上下文中使用，因为这些函数
		将禁用中断，而 ”napi“ 前缀的函数只可以在softirq上下文中使用。

drivers/net/ethernet/intel/i40e/i40e_txrx.c

+1 −3

Original line number	Diff line number	Diff line
		@@ -2144,9 +2144,7 @@ static struct sk_buff i40e_construct_skb(struct i40e_ring rx_ring,
		*/

		/* allocate a skb to store the frags */
		skb = __napi_alloc_skb(&rx_ring->q_vector->napi,
		I40E_RX_HDR_SIZE,
		GFP_ATOMIC \| __GFP_NOWARN);
		skb = napi_alloc_skb(&rx_ring->q_vector->napi, I40E_RX_HDR_SIZE);
		if (unlikely(!skb))
		return NULL;

drivers/net/ethernet/intel/i40e/i40e_xsk.c

+1 −2

Original line number	Diff line number	Diff line
		@@ -301,8 +301,7 @@ static struct sk_buff i40e_construct_skb_zc(struct i40e_ring rx_ring,
		net_prefetch(xdp->data_meta);

		/* allocate a skb to store the frags */
		skb = __napi_alloc_skb(&rx_ring->q_vector->napi, totalsize,
		GFP_ATOMIC \| __GFP_NOWARN);
		skb = napi_alloc_skb(&rx_ring->q_vector->napi, totalsize);
		if (unlikely(!skb))
		goto out;

drivers/net/ethernet/intel/iavf/iavf_txrx.c

+1 −3

Original line number	Diff line number	Diff line
		@@ -1334,9 +1334,7 @@ static struct sk_buff iavf_construct_skb(struct iavf_ring rx_ring,
		net_prefetch(va);

		/* allocate a skb to store the frags */
		skb = __napi_alloc_skb(&rx_ring->q_vector->napi,
		IAVF_RX_HDR_SIZE,
		GFP_ATOMIC \| __GFP_NOWARN);
		skb = napi_alloc_skb(&rx_ring->q_vector->napi, IAVF_RX_HDR_SIZE);
		if (unlikely(!skb))
		return NULL;