Let $P=\{1,\dots,n\}$ be a set of parties with local Hilbert spaces $\mathcal H^{(a)}$ of finite dimensions $d_a$. Following the notation of the local-invariant correlation-sector decomposition introduced by Aschauer \emph{et al.}\cite{aschauer}, choose for each party $a$ traceless Hermitian generators
\begin{equation*}
\sigma^{(a)}_1,\dots,\sigma^{(a)}_{d_a^2-1}
\end{equation*}
together with the identity $\sigma_0^{(a)}=\id$, normalized by
for $i,j\in\{0,1,\dots,d_a^2-1\}$, so that the orthogonality relation already includes the identity generator. For each nonempty subset $S\subseteq P$ we denote by
They are local-unitary invariants and admit a clean convexity-based entanglement criterion, but they retain only the total quadratic size of each tensor block.
The point of the present note is that the old geometric picture has a natural multipartite continuation. Instead of compressing each block to one scalar, we keep the full family of source-indexed correlation-response maps produced by a chosen source party and stack the responses landing in the different orthogonal sectors of the complement. In this sense the main contribution is best viewed as a new object architecture rather than a new functional applied to a familiar full correlation tensor. The general construction and cut-separable bound below work for arbitrary finite local dimensions; the later biseparable benchmark and all numerical examples then specialize to qubits, where the constants and geometry are especially transparent.
A substantial literature already studies separability criteria in the Bloch-representation and correlation-tensor language, including bipartite correlation-matrix criteria \cite{devicente,chenwu}, multipartite unfoldings and matricizations of the full correlation tensor \cite{hassanjoag,devicentehuber,li2014,jingzhang2023}, nonlinear geometric tensor criteria \cite{laskowski2011}, scalar multi-sector norm criteria \cite{klocklhuber2015}, and more recent extended or partition-adapted mixed-order block constructions \cite{shen2016,sarbicki2020,zhao2020,huang2024extended,liyaoyangfei2025}. This note is intended to sit explicitly within that landscape rather than outside it.
Historically, it is useful to distinguish two nearby lines of work. The local-invariant sector decomposition of Aschauer \emph{et al.}\cite{aschauer} already gave an early multipartite entanglement criterion in terms of the coefficients of a local operator expansion; that expansion is carried out, from the first equations of that paper onward, in the full product basis of local generators together with the identity on each party, so the full tensor $\mathcal C(\rho)$ used in Section~\ref{sec:tensor-viewpoint} below is already present there, and the sector-restricted tensor $C_S(\rho)$ used for $L_S$ is obtained from it by exactly the restriction to nonidentity indices already carried out in that paper. Later work such as Hassan and Joag \cite{hassanjoag} made the Bloch-representation terminology explicit and developed criteria from full-tensor unfoldings. The present construction belongs to that broader correlation-tensor lineage, but its organization is closest in spirit to, and its starting tensor is literally the one already used in, \cite{aschauer}.
It is worth stating carefully what is and is not being claimed here. The recent correlation-tensor literature already contains powerful criteria based on full tensors, matricizations, Bloch tensors, and cut-aware mixed-order block trace norms. In particular, recent generalized-Bloch constructions already study multipartite cut-aware mixed-order block trace-norm criteria \cite{liyaoyangfei2025}. Thus the intended novelty claim is narrow: not the first multipartite block trace-norm criterion of this general kind, but the specific direct-sum organization obtained by fixing one source party $a$ and stacking
Simultaneous bigraduation of the same cut matrix by \emph{both} source sector $V$ and target sector $T$, rather than one collapsed block & Definition~\ref{def:bigraduated}\\
Sub-block and sector-profile witnesses that this bigraduation makes available automatically, with no separate proof & Corollary~\ref{cor:sub-block}, profiles $\Phi_{a\to T}$, $\Phi^{(2)}_{a\to T}$\\
Sector profiles are demonstrably \emph{not} invariant under unitaries acting collectively on the coarse complement, even though the full cut norm is & Proposition following Corollary~\ref{cor:projection-bound}, Bell-product/CNOT example \\
Every shadow map, at every graduation level, is one matricization of a single augmented Bloch tensor, so the $\le1$ bound is inherited rather than reproved at each level & Theorem~\ref{thm:one-fact}, Section~\ref{sec:tensor-viewpoint}\\
An explicit, fully proved three-qubit biseparable threshold in closed form & Eq.~\eqref{eq:bisep-threshold}\\
A qutrit PPT-entangled benchmark and a systematic scan of $38$ four-qubit graph states and $n=3,4,5$ state families & qutrit benchmark in Section~\ref{sec:tensor-viewpoint}, Section~\ref{sec:qubit-numerics}\\
\caption{What this note recovers from the existing correlation-tensor literature (top) versus what it contributes on top of that baseline (bottom). The headline bound $\norm{\mathcal M_S(\rho)}_*\le1$ itself belongs to the top half; the substantive claims of the note are the structural and numerical items in the bottom half.
Two rigorously distinguished degeneracy mechanisms (continuous isotypic symmetry vs.\ discrete stabilizer support) explaining numerically observed singular-value degeneracies will be published separately \cite{aschauer2026b}.}
A guiding thread throughout the note is that this architecture is not tied to a single source party. Sections~\ref{sec:response-maps}--\ref{sec:multiparty-sources} show that fixing a source \emph{cluster}$S\subseteq P$ instead of a single party costs nothing in the proof: the rank-one mechanism behind the cut-separable bound is agnostic to how many parties sit on the source side. Section~\ref{sec:tensor-viewpoint} then makes precise in what sense this is not a coincidence: every construction in this note---the single-party map $\mathcal M_a$, its cluster generalization $\mathcal M_S$, and every sub-block compression of either---is a matricization or sub-block restriction of one and the same full correlation tensor, and the bound $\le1$ is a single rank-one fact about that tensor, inherited unchanged through each linear operation. The single-party case is simply the version of this fact that is cheapest to state first.
As a concrete demonstration that this stacked structure detects entanglement invisible to the standard PPT test, Section~\ref{sec:tensor-viewpoint} exhibits a two-qutrit state built from the Tiles unextendible product basis of \cite{bennettUPB}: it is numerically PPT to machine precision, so the Peres--Horodecki criterion is silent on it \cite{peres,horodeckiPPT}, yet the shadow-map witness detects its entanglement outright. We flag this example here because it is, in our view, the single clearest piece of evidence in the note that the construction has practical bite beyond reorganizing known bounds.
\subsection*{Notation guide}
The construction accumulates several closely related maps and scalars as it is refined step by step; Table~\ref{tab:notation} collects the main ones for reference, in the order they are introduced.
\begin{table}[h]
\centering
\small
\begin{tabular}{@{}lp{0.72\textwidth}@{}}
\toprule
Symbol & Meaning \\
\midrule
$C_S(\rho)$, $L_S(\rho)$& sector correlation tensor and its squared Frobenius norm (Aschauer \emph{et al.}\cite{aschauer}) \\
$\widetilde{\mathcal M}_a(\rho)$& intrinsic (basis-free) one-vs-rest response map, source party $a$\\
$\mathcal M_a(\rho)$& combined, normalized shadow map for a single source party $a$ (Definition~\ref{def:combined-shadow}) \\
$\widehat M_{a\to T}(\rho)$& sector shadow map, normalized independently of the other sectors \\
$\Phi_{a\to T}(\rho)$, $\Phi^{(2)}_{a\to T}(\rho)$& sector nuclear and Frobenius shadow profiles, $\norm{\widehat M_{a\to T}}_*$ and $\norm{\widehat M_{a\to T}}_\fro^2$\\
$\Phi_a^{(\le k)}(\rho)$, $\Phi_a^{(\ge k)}(\rho)$& nuclear norm after projecting onto grouped target sectors of order $\le k$ or $\ge k$\\
$\mathcal M_S(\rho)$& bigraduated shadow map for a source \emph{cluster}$S$ (Definition~\ref{def:bigraduated}); reduces to $\mathcal M_a(\rho)$ when $S=\{a\}$\\
$M_{V\to T}(\rho)$& bigraduated block from source sector $V\subseteq S$ to target sector $T\subseteq S^c$\\
$\Phi_{\mathrm{sym}}(\rho)$, $\Phi_{\max}(\rho)$& source-aggregated functionals, average and maximum of $\norm{\mathcal M_a(\rho)}_*$ over $a\in P$\\
$\mathcal C(\rho)$& full order-$n$ Bloch tensor of which every map above is a matricization or slice (Definition~\ref{def:full-tensor}) \\
$A_\lambda$& reduced shadow map on the multiplicity space of isotype $\lambda$, once $\rho$ carries a compatible symmetry (to be published in \cite{aschauer2026b}) \\
\caption{Main notation, in order of introduction. $\Phi_{\mathrm{sym}}$, $\Phi_{\max}$, and the biseparable threshold are qubit-specific; everything above them in the table is defined for arbitrary finite local dimensions.}
Fix a party $a\in P$, and write $\bar a=P\setminus\{a\}$. For each party $b$, let $\V^{(b)}$ be the real Hilbert space of Hermitian operators on $\mathcal H^{(b)}$, equipped with the Hilbert-Schmidt inner product, and let
After choosing orthonormal bases in $\V_0^{(a)}$ and in each $\V_T^{(\bar a)}$, these become the coordinate maps used below. We will use the shorter term \emph{shadow maps} for this family. We emphasize that this usage is unrelated to the classical-shadows measurement protocols of \cite{huangkuengpreskill2020}: the shadow maps of this note are linear response operators built directly from the correlation tensor of $\rho$, not estimators reconstructed from randomized single-copy measurements.
Under the orthogonal decomposition in Eq.~\eqref{eq:complement-bloch-decomposition}, this is just the matrix representation of the intrinsically defined map $\widetilde{\mathcal M}_a(\rho)$ in sector-adapted orthonormal coordinates. We refer to this direct-sum response operator as the \emph{combined shadow map}. The image of the unit ball in $\R^{d_a^2-1}$ under $\mathcal M_a(\rho)$ is the corresponding response ellipsoid in correlation space, which we also call the combined shadow ellipsoid of the party $a$.
\end{definition}
For qubits this reduces to the earlier normalization, since $(d_a-1)(d_{\bar a}-1)=2^{n-1}-1$. In particular, for three qubits this direct sum is
\begin{equation*}
\mathcal W_A=\R^3\oplus\R^3\oplus\R^9,
\end{equation*}
corresponding to the $B$, $C$, and $BC$ response sectors.
\subsection*{The cut-separable bound and its refinements}
The key fact is easiest to see first for states that are product across $a\mid\bar a$: then the full response operator is rank one, and the sector maps are simply its orthogonal components. The general cut-separable case follows by convexity.
\begin{theorem}
\label{thm:cut-bound}
Let $\rho$ be separable across the cut $a\mid\bar a$. Then
\begin{equation}
\norm{\mathcal M_a(\rho)}_*\le 1.
\label{eq:cut-bound}
\end{equation}
Consequently,
\begin{equation*}
\norm{\mathcal M_a(\rho)}_*>1
\qquad\Longrightarrow\qquad
\rho\text{ is entangled across }a\mid\bar a.
\end{equation*}
\end{theorem}
\begin{proof}
It is enough to begin with a product state across the cut,
\begin{equation*}
\rho=\rho_a\otimes\sigma_{\bar a}.
\end{equation*}
Let $r^{(a)}\in\V_0^{(a)}$ be the Bloch vector of $\rho_a$, defined by
\begin{equation*}
\langle X,r^{(a)}\rangle=\tr(\rho_a X),
\qquad X\in\V_0^{(a)},
\end{equation*}
and let $v_{\bar a}\in\V_0^{(\bar a)}$ be the traceless Bloch vector of $\sigma_{\bar a}$, defined analogously. Then
is rank one, and in sector-adapted coordinates this becomes the direct sum of the component maps. Equivalently, for each nonempty $T\subseteq\bar a$ let
using the correlation-sum identity from the Aschauer \emph{et al.} framework for the $(n-1)$-party state $\sigma_{\bar a}$. Thus Eq.~\eqref{eq:cut-bound} holds for every product state across the cut.
Now let $\rho$ be mixed and separable across the cut,
Let $\Pi$ be any orthogonal projection on $\V_0^{(\bar a)}$, and let $\Pi\mathcal M_a(\rho)$ denote the corresponding projected map in any orthonormal coordinates adapted to the decomposition of $\V_0^{(\bar a)}$. If $\rho$ is separable across the cut $a\mid\bar a$, then
\begin{equation}
\norm{\Pi\mathcal M_a(\rho)}_*\le 1.
\label{eq:projection-bound}
\end{equation}
Consequently, every orthogonally selected target subspace of $\V_0^{(\bar a)}$ yields a valid cut witness.
\end{corollary}
\begin{proof}
Orthogonal projection is contractive for the operator norm and hence for singular values. Therefore
The claim follows from Theorem~\ref{thm:cut-bound}.
\end{proof}
\begin{remark}
Equation~\eqref{eq:projection-bound} produces a whole hierarchy of weaker but natural cut witnesses. Besides the individual sectors $\Pi=P_T$, one can project onto grouped sector subspaces. For instance, for $1\le k\le |\bar a|$ let
Now $C_T(\sigma_{\bar a})$ depends only on the reduced state $\sigma_T$, so
\begin{equation*}
\norm{v_T}^2 = L_T(\sigma_T)
\le
\sum_{\emptyset\neq U\subseteq T}L_U(\sigma_T)
= d_T\tr(\sigma_T^2)-1
\le d_T-1.
\end{equation*}
Together with $\norm{r^{(a)}}^2\le d_a-1$, this proves the product-state case. The mixed separable case again follows by linearity and convexity, and the Frobenius statement follows from $\norm{X}_{\fro}\le\norm{X}_*$.
\end{proof}
\begin{definition}
For each nonempty subset $T\subseteq\bar a$, define the sector nuclear shadow profile by
which depends not only on the sizes of the individual sector maps but also on how their right-singular directions align in the common source space. Thus the full nuclear-norm signal is not, in general, a linear combination of the individual sector nuclear norms. In the cut-product case all sector maps share one common right factor, so the full map is again rank one and the proof of Theorem~\ref{thm:cut-bound} reduces to one Euclidean bound on the stacked target vector.
\end{remark}
\subsection*{Sector profiles are not cut invariants}
The sector profiles $\Phi_{a\to T}(\rho)$ and $\Phi^{(2)}_{a\to T}(\rho)$ are therefore not invariants of the coarse cut $a\mid\bar a$ under general collective complement unitaries; they are invariants only under unitaries that preserve the chosen internal factorization of $\bar a$ into parties.
\end{proposition}
\begin{proof}
Choose orthonormal bases of traceless Hermitian operators on $\mathcal H^{(a)}$ and $\mathcal H^{(\bar a)}$, normalized by
\begin{equation*}
\tr(\sigma_i\sigma_j)=d_a\,\delta_{ij},
\qquad
\tr(\tau_i\tau_j)=d_{\bar a}\,\delta_{ij}.
\end{equation*}
The intrinsic response map transforms by the adjoint actions on source and target Bloch spaces:
In the chosen orthonormal bases these adjoint actions are represented by real orthogonal matrices $O_a(U_a)$ and $O_{\bar a}(U_{\bar a})$, giving Eq.~\eqref{eq:two-sided-collective-covariance}. Left and right multiplication by orthogonal matrices preserve both nuclear and Frobenius norms, so the norm equalities follow.
The sector maps arise only after choosing the product operator basis on $\mathcal H^{(\bar a)}$ determined by the internal decomposition of $\bar a$ into parties and then splitting that basis into orthogonal summands. A general collective unitary on $\bar a$ need not preserve those summands, so it reshuffles the sector profile even though the full cut norm is unchanged.
\end{proof}
\begin{remark}
The failure of sector invariance is already visible in the simplest three-qubit Bell-product example. Let
So a collective unitary on $BC$ does not push the signal purely into the highest-order sector. Instead it redistributes a strongly localized $A\to B$ witness into a mixed profile spread across $A\to B$, $A\to C$, and $A\to BC$, while leaving the full $A\mid BC$ witness unchanged.
This persists under white noise. For
\begin{equation*}
\rho_p=p\rho+(1-p)\frac{\id}{8},
\qquad
\rho'_p=p\rho'+(1-p)\frac{\id}{8},
\end{equation*}
the full-map threshold is the same in both forms,
\begin{equation*}
\norm{\mathcal M_A(\rho_p)}_*>1
\iff
\norm{\mathcal M_A(\rho'_p)}_*>1
\iff
p>\frac{1}{\sqrt 6}.
\end{equation*}
But the sector thresholds differ sharply: before the collective rotation the sector $A\to B$ already detects for $p>1/3$, whereas after the rotation the sectors $A\to B$ and $A\to C$ never strictly violate the cut-separable bound and the sector $A\to BC$ only detects for $p>\sqrt{3/8}$. In particular, at $p=0.60$ one has $\norm{\mathcal M_A(\rho'_p)}_*>1$ while all three individual sectors still satisfy $\Phi_{A\to T}(\rho'_p)\le1$.
The construction of Section~\ref{sec:response-maps} singles out one party $a$ as source and treats the entire complement $\bar a$ as target. Nothing in the argument in fact requires $|S|=1$ on the source side; the same object exists for any source cluster $\emptyset\neq S\subsetneq P$, with complement $S^c:=P\setminus S$. Making this explicit exposes a layer of internal structure on the source side that Definition~\ref{def:combined-shadow} discards by construction, and it costs nothing beyond re-reading the definitions and the proof of Theorem~\ref{thm:cut-bound} with $a$ replaced by $S$.
\subsection*{Source sector decomposition}
For nonempty $S\subseteq P$, set $\V^{(S)}:=\bigotimes_{a\in S}\V^{(a)}$ and apply exactly the decomposition of Eq.~\eqref{eq:complement-bloch-decomposition}, now with $S$ in the role previously played by $\bar a$:
This is not a new construction, only the source-side instance of the same orthogonal sector decomposition already used for the complement. The traceless source space is accordingly
graded by which parties within $S$ are active. For $S=\{a\}$ the only nonempty $V\subseteq S$ is $V=S$ itself, so this decomposition is trivial for a singleton source; the bigraduation below is genuinely new structure only once $|S|\ge2$.
is given by the same formula as Eq.~\eqref{eq:intrinsic-response}, with $a$ replaced by $S$. Both $\V_0^{(S)}$ and $\V_0^{(S^c)}$ carry an orthogonal sector decomposition, Eq.~\eqref{eq:source-bloch-decomposition} on the source side and Eq.~\eqref{eq:complement-bloch-decomposition} on the target side, so the coordinate representation of $\widetilde{\mathcal M}_S(\rho)$ is naturally \emph{bigraded} by source sector $V\subseteq S$ and target sector $T\subseteq S^c$ simultaneously. Writing $\iota_V:\V_V^{(S)}\hookrightarrow\V_0^{(S)}$ for the canonical inclusion of a source sector and $P_T:\V_0^{(S^c)}\to\V_T^{(S^c)}$ for the orthogonal projection onto a target sector, define
i.e.\ the matrix representation of $\widetilde{\mathcal M}_S(\rho)$, normalized exactly as in Eq.~\eqref{eq:combined-map}, in sector-adapted orthonormal coordinates on both sides.
\end{definition}
For $S=\{a\}$, Definition~\ref{def:bigraduated} reduces exactly to Definition~\ref{def:combined-shadow}: the source-side decomposition then has only the single summand $V=S$, so the bigraduation collapses to the ordinary target-only graduation of $\mathcal M_a$. Definition~\ref{def:combined-shadow} is thus the singleton case of this construction, not a separate object introduced in parallel to it.
The total number of scalar entries in $\mathcal M_S(\rho)$ is $(d_S-1)(d_{S^c}-1)$, exactly as for the collapsed cut matrix, so evaluating the norm $\norm{\mathcal M_S(\rho)}_*$ in Theorem~\ref{thm:cluster-cut} costs a single singular value decomposition of that size and is no more expensive than the realignment-type bounds it recovers. The bigraduation of Definition~\ref{def:bigraduated} does not change this cost; it only reorganizes the same entries into $(2^{|S|}-1)(2^{|S^c|}-1)$ combinatorial blocks $M_{V\to T}$. Consequently, using the full map $\mathcal M_S(\rho)$ as a single witness remains cheap, but exploiting the sub-block witnesses of Corollary~\ref{cor:sub-block} exhaustively --- inspecting every source sector against every target sector separately, rather than only the handful used in the examples below --- requires examining up to $O(2^{|S|+|S^c|})=O(2^n)$ individual blocks in the worst case. We will show in \cite{aschauer2026b} that when $\rho$ carries a compatible symmetry, this exponential proliferation of combinatorial blocks collapses instead onto a typically much smaller number of representation-theoretic multiplicity spaces, giving a tractable route to the same sub-block information without enumerating every sector by hand.
Let $\rho$ be separable across the cut $S\mid S^c$. Then
\begin{equation}
\norm{\mathcal M_S(\rho)}_*\le1.
\label{eq:cluster-cut-bound}
\end{equation}
Consequently, $\norm{\mathcal M_S(\rho)}_*>1$ implies that $\rho$ is entangled across $S\mid S^c$.
\end{theorem}
\begin{proof}
The proof of Theorem~\ref{thm:cut-bound} goes through with $a\to S$ and $\bar a\to S^c$ without modification. For a product state $\rho=\rho_S\otimes\sigma_{S^c}$, let $r^{(S)}\in\V_0^{(S)}$ be the traceless Bloch vector of the $|S|$-party reduced state $\rho_S$ -- the same object as $r^{(a)}$ in the proof of Theorem~\ref{thm:cut-bound}, now for the composite system $S$ treated as a single $d_S$-dimensional party -- and let $v_{S^c}\in\V_0^{(S^c)}$ be defined analogously from $\sigma_{S^c}$. Tensor factorization gives $\widetilde{\mathcal M}_S(\rho)=v_{S^c}(r^{(S)})^T$, which is rank one, and
by the same correlation-sum identity used in the proof of Theorem~\ref{thm:cut-bound}, now applied to the $|S|$-party state $\rho_S$ and the $|S^c|$-party state $\sigma_{S^c}$ rather than to single-party marginals. Hence $\norm{\mathcal M_S(\rho)}_*\le1$ for every product state across the cut, and convexity of the nuclear norm extends the bound to mixtures exactly as in the proof of Theorem~\ref{thm:cut-bound}.
\end{proof}
\begin{corollary}[Sub-block witnesses]
\label{cor:sub-block}
Let $\mathcal V\subseteq\{V:\emptyset\neq V\subseteq S\}$ and $\mathcal T\subseteq\{T:\emptyset\neq T\subseteq S^c\}$ be any nonempty families, and let $P_{\mathcal V}$, $P_{\mathcal T}$ be the orthogonal projections onto $\bigoplus_{V\in\mathcal V}\V_V^{(S)}$ and $\bigoplus_{T\in\mathcal T}\V_T^{(S^c)}$ respectively. If $\rho$ is separable across $S\mid S^c$, then
Orthogonal projections are contractions for the operator norm, and the nuclear norm satisfies $\norm{AXB}_*\le\norm{A}_{\mathrm{op}}\norm{X}_*\norm{B}_{\mathrm{op}}$ for linear maps $A,B$ of compatible size. Taking $A=P_{\mathcal T}$ and $B=P_{\mathcal V}$ gives Eq.~\eqref{eq:sub-block-bound}; the claim then follows from Theorem~\ref{thm:cluster-cut}. This is the two-sided extension of Corollary~\ref{cor:projection-bound}, which only ever compressed the target side.
\end{proof}
In particular, taking $\mathcal V=\{S\}$ isolates the single block $M_{S\to T}$, which carries only the correlation attributable to the full cluster $S$ acting jointly rather than to any proper sub-cluster of $S$ -- a witness targeted specifically at ``genuine $S$'' structure landing in the sector $T$.
\subsection*{Example: the Smolin state under two different cuts}
where here and in what follows for qubits $X,Y,Z$ denote the Pauli matrices, i.e.\ $\sigma_1=X$, $\sigma_2=Y$, $\sigma_3=Z$ in the generator convention fixed formally in Section~\ref{sec:tensor-viewpoint} below. We evaluate it under two cuts side by side.
\emph{The $2\mid2$ cut.} Take the cluster $S=\{A,B\}$, $S^c=\{C,D\}$. Every block $M_{V\to T}$ with $V\neq S$ or $T\neq S^c$ vanishes identically, because the Smolin state carries no one- or three-body correlations. Only $M_{S\to S^c}$ survives, as the $9\times9$ matrix diagonal on the aligned Pauli directions $(x,x)\to(x,x)$, $(y,y)\to(y,y)$, $(z,z)\to(z,z)$ with unit coefficients and zero elsewhere. After the normalization $1/\sqrt{(d_S-1)(d_{S^c}-1)}=1/3$ this gives three singular values of $1/3$ each, so
\begin{equation*}
\norm{\mathcal M_{AB}(\rho_{\mathrm{Smo}})}_*=1
\quad\text{exactly --- saturating, not violating, the bound of Theorem~\ref{thm:cluster-cut}.}
\end{equation*}
\emph{The $1\mid3$ cut.} Take instead a single source party $a$, so $\bar a$ is the remaining three-qubit cluster. Again only the full four-body sector survives, giving three orthogonal response directions with unnormalized singular value $1$ each. The normalization is now $1/\sqrt{(d_a-1)(d_{\bar a}-1)}=1/\sqrt{1\cdot7}=1/\sqrt7$, so
strictly violating the bound of Theorem~\ref{thm:cut-bound}.
So the same state sits exactly at the boundary for every $2\mid2$ cut while clearly violating the $1\mid3$ bound --- the two cut types are correctly told apart within one framework. The refinement from singleton to cluster sources costs nothing in the proof yet makes this distinction available at all, since the singleton construction of Definition~\ref{def:combined-shadow} cannot even pose the $2\mid2$ question. Section~\ref{sec:qubit-numerics} below returns to the $1\mid3$ value in the source-aggregated language of $\Phi_{\mathrm{sym}}$ and $\Phi_{\max}$, and adds the white-noise robustness threshold.
%% \subsection*{Example: the Smolin state under a $2\mid2$ cut}
%% (examined again from the single-party viewpoint in Section~\ref{sec:qubit-numerics} below), and take the cluster $S=\{A,B\}$, $S^c=\{C,D\}$. Every block $M_{V\to T}$ with $V\neq S$ or $T\neq S^c$ vanishes identically, because the Smolin state carries no one- or three-body correlations. Only $M_{S\to S^c}$ survives, as the $9\times9$ matrix diagonal on the aligned Pauli directions $(x,x)\to(x,x)$, $(y,y)\to(y,y)$, $(z,z)\to(z,z)$ with unit coefficients and zero elsewhere. After the normalization $1/\sqrt{(d_S-1)(d_{S^c}-1)}=1/3$ this gives three singular values of $1/3$ each, so
%% \quad\text{exactly --- saturating, not violating, the bound of Theorem~\ref{thm:cluster-cut}.}
%% \end{equation*}
%% This is consistent with the Smolin state being separable across every $2\mid2$ cut while violating the $1\mid3$ bound at $3/\sqrt7\approx1.134$, computed below in Section~\ref{sec:qubit-numerics}. The refinement from singleton to cluster sources costs nothing in the proof yet correctly distinguishes the two cut types, where the singleton construction of Definition~\ref{def:combined-shadow} cannot even pose the $2\mid2$ question.
\section{The tensor viewpoint: shadow maps as unfoldings of one full Bloch tensor}
\label{sec:tensor-viewpoint}
The guiding thread announced in Section~\ref{sec:response-maps}--\ref{sec:multiparty-sources} can now be made precise. Every non-scalar object introduced so far --- $M_a$, its bigraduated extension $\mathcal M_S$, and the sub-block witnesses of Corollary~\ref{cor:sub-block} --- turns out to be a matricization or sub-block restriction of a single order-$n$ tensor built from the full correlation data of $\rho$; the source-aggregated functionals used in the qubit applications below are then averages or maxima of the corresponding one-vs-rest norms. The $\le1$ bound is therefore not a family of independently proved facts, but one rank-one statement about that tensor, observed through different linear lenses.
\begin{definition}[Full Bloch tensor]
\label{def:full-tensor}
For each party $a$, we augment the local index range by including $i_a =0$, which selects the identity $\sigma_0^{(a)}=\id$ already introduced in Section~\ref{sec:old-criterion}. This allows us to define the full (order-$n$) Bloch tensor
Since $\bigotimes_{a\in P}\R^{d_a^2}=\bigotimes_{a\in P}\bigl(\R\oplus\R^{d_a^2-1}\bigr)$ expands by distributivity into $2^n$ orthogonal summands indexed by which legs are trivial, every sector tensor of Eq.~\eqref{eq:corr-tensor-def} is simply a slice of $\mathcal C(\rho)$:
\begin{equation}
C_V(\rho)=\mathcal C(\rho)\big|_{\,i_a\neq0\text{ for }a\in V,\ i_a=0\text{ for }a\notin V},
The sector decomposition used throughout this note, on both the target side (Section~\ref{sec:response-maps}) and the source side (Section~\ref{sec:multiparty-sources}), is therefore not an additional structure imposed on the correlation data: it \emph{is} the tensor-product structure of $\mathcal C(\rho)$ in the identity-plus-generators basis.
\begin{remark}[This is already the tomography tensor of \cite{aschauer}]
\label{rem:aschauer-tensor}
Definition~\ref{def:full-tensor} introduces no object beyond what \cite{aschauer} starts from. Writing out the operator expansion of $\rho$ in the full product basis $\{\bigotimes_{a\in P}\sigma^{(a)}_{i_a}:0\le i_a\le d_a^2-1\}$ used there for state tomography gives exactly
with $\mathcal C(\rho)=(c_{i_1,\dots,i_n})$, over the same unrestricted index range. The sector-restricted tensor $C_S(\rho)$ of Eq.~\eqref{eq:corr-tensor-def}, on which the correlation strengths $L_S$ and every construction built on them in this note ultimately depend, is the further restriction of that same tomography tensor to $i_a>0$ for every $a\in S$. In this sense, Sections~\ref{sec:response-maps}--\ref{sec:multiparty-sources} never leave the object \cite{aschauer} already had in hand; what changes is only what is extracted from it, replacing the scalar sector norm $L_S$ with a matrix unfolding and its singular values. The point of the present section is that this change of extraction is itself best understood at the level of the tensor $\mathcal C(\rho)$, rather than sector by sector.
\end{remark}
We now switch perspectives: the shadow maps of the preceding sections will no longer be treated as separately constructed response operators, but as canonical unfoldings, sector restrictions, and identity-leg slices of the single full tensor $\mathcal C(\rho)$.
\begin{proposition}[Product states are exactly the states whose full Bloch tensor has CP-rank one]
\label{prop:cp-rank-one}
$\rho=\bigotimes_{a\in P}\rho_a$ if and only if $\mathcal C(\rho)=\bigotimes_{a\in P}w^{(a)}$ for vectors $w^{(a)}\in\R^{d_a^2}$ with $w^{(a)}_0=1$.
\end{proposition}
\begin{proof}
($\Rightarrow$) Immediate from multiplicativity of the trace over the tensor factors, with $w^{(a)}_{i_a}:=\tr(\rho_a\sigma^{(a)}_{i_a})$.
($\Leftarrow$) By the orthogonality relation~\eqref{eq:generator-orthogonality}, extended over $i,j\in\{0,\dots,d_a^2-1\}$, the map $\rho\mapsto\mathcal C(\rho)$ is a linear bijection with inverse
Substituting a rank-one $\mathcal C(\rho)=\bigotimes_a w^{(a)}$ into this inversion formula factorizes term by term into $\bigotimes_{a\in P}\rho_a$ with $\rho_a:=d_a^{-1}\sum_{i_a}w^{(a)}_{i_a}\sigma^{(a)}_{i_a}$. The condition $w_0^{(a)}=1$ gives $\tr(\rho_a)=1$ for each $a$; positivity of each $\rho_a$ then follows because $\rho=\bigotimes_a\rho_a$ is positive semidefinite and every factor is Hermitian with trace one and nonzero.
\end{proof}
\begin{proposition}[Shadow maps are unfoldings]
\label{prop:unfolding}
Fix a bipartition $P=S\sqcup S^c$. Grouping the $S$-legs of $\mathcal C(\rho)$ into a single row index and the $S^c$-legs into a single column index is the standard mode-$(S,S^c)$ matricization of $\mathcal C(\rho)$ in the sense of the multilinear singular value decomposition \cite{delathauwer}. Deleting the trivial ($i=0$) row and column --- equivalently, discarding the $V=\emptyset$ and $T=\emptyset$ sectors, which carry no information beyond normalization --- and rescaling by $[(d_S-1)(d_{S^c}-1)]^{-1/2}$ reproduces $M_S(\rho)$ exactly. The bigraduated shadow map $\mathcal M_S(\rho)$ of Definition~\ref{def:bigraduated} is the same unfolding with the row index additionally kept graded by $V\subseteq S$ instead of collapsed.
Fix a bipartition $P=S\sqcup S^c$. For every state separable across $S\mid S^c$, the normalized nuclear-norm bound $\norm{\cdot}_*\le1$ holds not only for the full shadow map $\mathcal M_S(\rho)$, but also for every witness obtained from the same mode-$(S,S^c)$ unfolding of $\mathcal C(\rho)$ by keeping or collapsing the source and target sector gradings, by taking two-sided sector restrictions as in Corollary~\ref{cor:sub-block}, or by taking target-side identity-leg slices corresponding to partial traces as in Corollary~\ref{cor:trace-is-slice} below, with the normalization appropriate to the remaining source and target systems. Thus Theorem~\ref{thm:cut-bound}, its cluster generalization (Theorem~\ref{thm:cluster-cut}), and the sub-block witnesses of Corollary~\ref{cor:sub-block} are not independently proved facts, but one algebraic statement observed through different linear lenses.
\end{theorem}
\begin{proof}
It is enough to consider a product state across the chosen cut, $\rho=\rho_S\otimes\sigma_{S^c}$. Grouping the legs in $S$ and $S^c$, trace multiplicativity gives
so the mode-$(S,S^c)$ unfolding is rank one. After deleting the identity row and column, this is the rank-one matrix $v_{S^c}(r^{(S)})^T$ appearing in the proof of Theorem~\ref{thm:cluster-cut}, and the same correlation-sum identity gives
\begin{equation*}
\norm{r^{(S)}}^2\le d_S-1,
\qquad
\norm{v_{S^c}}^2\le d_{S^c}-1.
\end{equation*}
Hence the normalized full unfolding has nuclear norm at most one for each product term.
Keeping the sector gradings is only a change of coordinates, while collapsing them gives the same matrix representation with grouped row or column indices. Two-sided sector restrictions have the form $Axy^TB=(Ax)(B^Ty)^T$ on each rank-one product term and cannot increase the nuclear norm when $A$ and $B$ are orthogonal projections. Target-side identity-leg slices give the corresponding reduced product tensor and obey the same estimate with the dimensions of the surviving source and target systems. Finally, linearity of $\rho\mapsto\mathcal C(\rho)$ and convexity of the nuclear norm extend the bound from product states to arbitrary mixtures separable across $S\mid S^c$.
\end{proof}
\begin{corollary}[Partial trace is a slice, not a sum]
\label{cor:trace-is-slice}
For $E\subseteq P$ and $\rho_{P\setminus E}:=\tr_E(\rho)$,
\begin{equation}
\mathcal C(\rho_{P\setminus E})=\mathcal C(\rho)\big|_{\,i_a=0\text{ for all }a\in E},
\label{eq:trace-is-slice}
\end{equation}
i.e.\ the marginal's full tensor is the slice of $\mathcal C(\rho)$ at the trivial index on every traced-out leg, not a contraction or summation over $E$. Consequently, if $P=S\sqcup R\sqcup E$ with $S$ a source cluster as in Definition~\ref{def:bigraduated}, $\rho_{SR}:=\tr_E(\rho)$, and $\Pi_R$ denotes the orthogonal projection that annihilates every target sector $T$ with $T\cap E\neq\emptyset$, then exactly
Eq.~\eqref{eq:trace-is-slice} is immediate from $\sigma_0^{(a)}=\id$: setting $i_a=0$ for $a\in E$ in the defining sum of $\mathcal C(\rho)$ inserts the identity on every traced-out leg, which is exactly $\tr_E(\rho)$ evaluated against the remaining generators. For the second claim, apply Proposition~\ref{prop:unfolding} to the slice~\eqref{eq:trace-is-slice}: because $S\cap E=\emptyset$, the source legs are untouched by the slicing, so for every $T\subseteq R$ the unnormalized block $M_{S\to T}$ computed from $\mathcal C(\rho)$ agrees exactly with the one computed from $\mathcal C(\rho_{SR})$. The two combined maps therefore differ only in their normalization constants, $[(d_S-1)(d_{S^c}-1)]^{-1/2}$ for $\mathcal M_S(\rho)$ against $[(d_S-1)(d_R-1)]^{-1/2}$ for $\mathcal M_S(\rho_{SR})$, since the complement of $S$ is $R\cup E$ in the first case and $R$ alone in the second, with $d_{S^c}=d_Rd_E$. Their ratio is exactly the stated factor.
\end{proof}
Eq.~\eqref{eq:trace-rescale} shows that discarding a residual cluster $E$ by tracing it out is a strictly weaker operation than the sub-block compression of Corollary~\ref{cor:sub-block}: the latter only ever shrinks the nuclear norm, while Eq.~\eqref{eq:trace-rescale} rescales it upward by the factor $\sqrt{(d_{S^c}-1)/(d_R-1)}\ge1$, so that a violation of the bound on $\rho_{SR}$ can certify entanglement across $S\mid R$ that survives the complete loss of $E$, a strictly stronger and operationally different statement than merely detecting entanglement somewhere across $S\mid RE$.
\begin{remark}[Why unfold at all]
The full tensor $\mathcal C(\rho)$ carries strictly more information than any single unfolding: two states can share every matricization $M_S$ over all bipartitions and still differ in genuine multi-way structure, exactly as a generic tensor is not determined by its unfoldings alone. The reason this note works with unfoldings rather than $\mathcal C(\rho)$ directly is computational, not conceptual. The CP-rank-one statement for products, Proposition~\ref{prop:cp-rank-one}, is exact and dimension-independent, but the associated \emph{tensor} nuclear norm --- the natural generalization of $\norm{\cdot}_*$ that would witness separability directly on $\mathcal C(\rho)$, as the infimum of $\sum_r|\lambda_r|$ over CP decompositions --- has no polynomial-time algorithm once three or more legs are grouped independently. Every matricization used in this note, by contrast, is an ordinary matrix with a computable singular value decomposition. The constructions of the preceding sections are thus best understood as the maximal set of efficiently computable shadows of one underlying rank-one fact, chosen along the cut structure that is operationally relevant to entanglement questions.
\end{remark}
\subsection*{A qutrit PPT-entangled benchmark}
The dimension-independent normalization is not only a formal convenience. As a two-qutrit test case, consider the Tiles unextendible product basis of Bennett
\emph{et al.}~\cite{bennettUPB}, consisting of the five orthonormal product vectors
This is the standard rank-four PPT-entangled state supported on the completely entangled complement of the UPB. Using Gell-Mann generators scaled by $\sqrt{3/2}$, so that $\tr(\sigma_i\sigma_j)=3\delta_{ij}$ as in Eq.~\eqref{eq:generator-orthogonality}, the bipartite shadow map is the $8\times8$ correlation matrix divided by
Thus the state is PPT up to numerical precision, so the Peres-Horodecki PPT test is silent \cite{peres,horodeckiPPT}, while the shadow-map witness detects its entanglement. This should be read as complementarity rather than domination: the realignment criterion of Chen and Wu \cite{chenwu} also detects this benchmark, with trace norm $1.087412465\ldots$. The script \texttt{scripts/tiles\_upb.py} reproduces the generator normalization, the PPT spectrum, the shadow value, and the realignment comparison.
\subsection*{Qubit Pauli-tensor form}
For qubits we use the convention already anticipated in the Smolin example of Section~\ref{sec:multiparty-sources}: $\sigma_0=\id$ and $\sigma_1=X$, $\sigma_2=Y$, $\sigma_3=Z$. In this notation, the full Bloch tensor $\mathcal C(\rho)$ becomes the familiar Pauli-correlation tensor
and each sector is specified simply by the support pattern of the nonidentity Pauli indices. Thus the one-vs-rest map for a source party $a$ is obtained by unfolding this Pauli tensor with the $a$-leg as source, discarding the all-identity sector on the complement, and multiplying by
\begin{equation*}
\frac{1}{\sqrt{2^{n-1}-1}}.
\end{equation*}
More generally, for a qubit source cluster $S$ the normalized cluster map is the Pauli-tensor unfolding across $S\mid S^c$, with the all-identity source and target sectors removed, scaled by
\begin{equation*}
\frac{1}{\sqrt{(2^{|S|}-1)(2^{|S^c|}-1)}}.
\end{equation*}
This form makes two features of the examples below transparent. First, adding white noise as
\begin{equation*}
\rho(p)=p\rho+(1-p)\frac{\id}{2^n}
\end{equation*}
leaves the all-identity coefficient fixed and multiplies every nonidentity Pauli coefficient by $p$, so every shadow map and every shadow norm scales linearly with $p$. Second, stabilizer and graph states have Pauli tensors supported on their stabilizer groups, with nonzero coefficients equal to $\pm1$. Their shadow maps are therefore normalized signed support-pattern unfoldings, which explains why the numerical graph-state values below are rigid singular-value facts rather than generic floating-point coincidences.
\section{Qubit specialization and source-aggregated benchmarks}
For the remaining benchmarks and numerics we stay in the qubit setting. The single-party response spaces are $\R^3$, the normalization in Eq.~\eqref{eq:combined-map} reduces to $1/\sqrt{2^{n-1}-1}$, and one can derive explicit constants that do not seem to be available so cleanly in higher dimensions.
Hence violating either bound certifies entanglement.
\end{corollary}
\begin{proof}
A fully separable state is separable across every one-vs-rest cut. The claim follows by applying Theorem~\ref{thm:cut-bound} to each party.
\end{proof}
In equal local dimensions these functionals are permutation invariant. More generally, they are source-aggregated cut-sensitive scalars. The average probes all one-vs-rest cuts simultaneously, whereas the maximum asks whether at least one cut exhibits a large combined shadow ellipsoid.
For three qubits the symmetric average can be optimized explicitly over the biseparable set. This is the first place where we use genuinely qubit-specific formulas rather than only the general response-map architecture. Suppose first that
\begin{equation*}
\rho=\rho_A\otimes\sigma_{BC}
\end{equation*}
is product across $A\mid BC$. If the total state is pure, then this is equivalent to being pure and separable across that cut. Let $a\in\R^3$ be the Bloch vector of $\rho_A$, and let
\begin{equation*}
b,c\in\R^3,
\qquad
T\in\R^{3\times 3}
\end{equation*}
be the one- and two-body correlation data of the pure two-qubit state $\sigma_{BC}$. Then $\norm{a}=1$, and Theorem~\ref{thm:cut-bound} gives
the relevant response vectors for each source party are orthogonal, with equal squared norm $2/3$ after the normalization in Eq.~\eqref{eq:combined-map}. Hence the singular values of each $\mathcal M_a$ are all equal to $\sqrt{2/3}$, so
so the criterion correctly detects entanglement across every $1\mid3$ cut, while --- as already noted --- this is not a genuine-multipartite conclusion, since the Smolin state is separable across every $2\mid2$ split. For the white-noise family $p\rho_{\mathrm{Smo}}+(1-p)\id/16$, the fully separable threshold would only be crossed for
\begin{equation*}
p>\frac{\sqrt 7}{3}\approx 0.882,
\end{equation*}
so this example is much more fragile under white noise than the graph-state families discussed below.
Using the \texttt{qtensor} package, we also evaluated the symmetric shadow functionals for all connected labeled graph states on four qubits. Numerically, all $38$ such graph states give the same value,
The same numerical scan shows that this shadow detection is not explained by pairwise entanglement in the reduced states. For those representative families, every two-qubit marginal remains PPT at the threshold $p=\sqrt7/6$, and for the line and ring graph states some of the two-qubit marginals are even maximally mixed. So the shadow functional is responding to multipartite correlation structure that is not visible in pair reductions. This is still not a four-qubit genuine-multipartite-entanglement proof, because the corresponding biseparable threshold is not yet known, but it makes the criterion promising as a genuinely multipartite diagnostic.
The ring graph state also gives a compact illustration of what the bigraduated $2\mid2$ map sees beyond the source-aggregated scalars. Let $\rho_{\square}$ be the four-qubit graph state on the cycle with edges $(1,2),(2,3),(3,4),(4,1)$. Evaluating the normalized cluster map of Eq.~\eqref{eq:bigraduated-map} across the two inequivalent $2\mid2$ cuts gives
\begin{center}
\begin{tabular}{lccc}
\toprule
cut $S\mid S^c$&$\norm{\mathcal M_S(\rho_{\square})}_*$&$\norm{P_{S^c}\mathcal M_S(\rho_{\square})P_S}_*$& marginal on $S$\\
Here $P_S$ and $P_{S^c}$ denote the projections onto the full source and full target sectors, so the middle column isolates the genuine two-body-to-two-body block within the same normalized witness. Thus both cuts violate the separable bound, but by different mechanisms: for the adjacent cut the full-sector block already violates, while for the diagonal cut that block only saturates the bound and the excess comes from lower source or target sectors. The script \texttt{scripts/grraph\_state\_cuts.py} reproduces these values directly from the Pauli correlation tensor.
As a first systematic extension, we also scanned the families $\GHZ_n$, $W_n$, the line graph state, and the ring graph state for $n=3,4,5$, together with D\"ur states for $n=4,5$. Numerically,
throughout that range, with common values $\sqrt6$, $6/\sqrt7$, and $2.19089\ldots$ for $n=3,4,5$, respectively. The $W_n$ family is consistently slightly lower but still well above the fully separable bound, while the D\"ur family already lies below $1$ for $n=4,5$. So the symmetric shadow functional strongly favors graph-like and GHZ-like global correlation structure, but it is not simply a monotone of party number.
At $n=4$, for instance, $W_4$ still crosses the fully separable white-noise threshold at about $p\approx0.469$, whereas the D\"ur value is already below the fully separable benchmark even before white noise is added.
To probe the unresolved four-qubit biseparable benchmark, we performed a small random search over pure biseparable states across all inequivalent cuts, using $120$ samples per cut. The largest sampled value was
\begin{equation*}
\Phi_{\mathrm{sym}}\approx 2.235
\end{equation*}
for a state separable across a $2\mid2$ partition, while the best sampled $1\mid3$ values were only around $1.94$. This is not a proof of the true biseparable threshold, but it suggests two useful heuristics: first, the most dangerous competitors to the graph-state value $6/\sqrt7\approx2.268$ come from $2\mid2$ cuts rather than $1\mid3$ cuts; second, the connected four-qubit graph-state value sits slightly above the best random biseparable samples we found.
The combined shadow map should be viewed as a structured refinement of the older correlation strengths $L_S$ introduced by Aschauer \emph{et al.}\cite{aschauer}. Those quantities keep one Frobenius norm per tensor block; the present construction keeps the common one-vs-rest channel structure across all orthogonal sectors on the complement. In that sense it preserves the geometric spirit of that local-invariant sector decomposition while extracting more information from the same correlation data.
First, the sharp multipartite threshold beyond the three-qubit case. While the symmetric average admits a closed-form optimization over the three-qubit biseparable set, the corresponding four-qubit threshold remains unknown. Our numerical search suggests that the extremal competitors arise from $2\mid2$ partitions rather than $1\mid3$ cuts, providing a concrete starting point for a future analytic treatment.
Second, the remaining residual-cluster construction. Of the three options for a cluster $E$ belonging to neither the source $S$ nor the target $R$, the first two --- folding $E$ into the target side and tracing it out --- are now fully covered: the former is the construction used throughout this note, and the latter is exactly Corollary~\ref{cor:trace-is-slice}, which shows it to be a rescaled restriction of the same object certifying a strictly stronger, loss-of-$E$-robust form of $S\mid R$ entanglement. The third option, keeping $E$ as its own graded tensor leg alongside $S$ and $R$ instead of folding or tracing it out, leaves matrix territory entirely for the genuine higher-order tensor discussed in the closing remark of Section~\ref{sec:tensor-viewpoint}, with the computability cost noted there. A systematic treatment of this third, genuinely tensorial option is left for future work.
Third, there is a natural exact continuation of this program at the level of moment problems. In the qubit case, the full Pauli correlation coefficients determine the density matrix linearly, so the separability problem can be formulated as a truncated moment problem on a product of Bloch spheres, in line with recent moment-based tensor criteria \cite{huang2024moments}. In higher local dimensions the same philosophy should run through generalized Bloch coordinates and products of local state spaces, with dimension-dependent response spaces and cut bounds. From that viewpoint the present shadow criteria are inexpensive front-end tests: they keep enough geometry to matter, but remain explicit and analytically tractable. The full moment hierarchy belongs to a larger project, so in the present note we retain it only as an outlook.