Abstract:In this paper, we prove topology dependent bounds on the number of rounds needed to compute Functional Aggregate Queries ($\FAQ$s) studied by Abo Khamis et al. [PODS 2016] in a synchronous distributed network under the model considered by Chattopadhyay et al. [FOCS 2014, SODA 2017]. Unlike the recent work on computing database queries in the Massively Parallel Computation model, in the model of Chattopadhyay et al., nodes can communicate only via private point-to-point channels and we are interested in bounds that work over an \em arbitrary communication topology. This model, which is closer to the well-studied $\congest$ model in distributed computing and generalizes Yao's two party communication complexity model, has so far only been studied for problems that are common in the two-party communication complexity literature. This is the first work to consider more practically motivated problems in this distributed model. For the sake of exposition, we focus on two specific problems in this paper: Boolean Conjunctive Query ($\BCQ$) and computing variable/factor marginals in Probabilistic Graphical Models (PGMs). We obtain tight bounds on the number of rounds needed to compute such queries as long as the underlying hypergraph of the query is $O(1)$-degenerate and has $O(1)$-arity. In particular, the $O(1)$-degeneracy condition covers most well-studied queries that are efficiently computable in the centralized computation model like queries with constant treewidth. These tight bounds depend on a new notion of 'width' (namely \em internal-node-width ) for Generalized Hypertree Decompositions (GHDs) of acyclic hypergraphs, which minimizes the number of internal nodes in a sub-class of GHDs. To the best of our knowledge, this width has not been studied explicitly in the theoretical database literature. Finally, we consider the problem of computing the product of a vector with a chain of matrices and prove tight bounds on its round complexity (over a finite field of two elements) using a novel min-entropy based argument.

Making problems tractable on big data via preprocessing with polylog-size output

Tractable Queries on Big Data Via Preprocessing with Logarithmic-Size Output

Sublinear-time Reductions for Big Data Computing

Tractable Circuits in Database Theory

Discovering Dichotomies for Problems in Database Theory

A Qubit, a Coin, and an Advice String Walk Into a Relational Problem

General Space-Time Tradeoffs via Relational Queries

Taming Quantum Time Complexity

What do Shannon-type Inequalities, Submodular Width, and Disjunctive Datalog have to do with one another?

Small hitting-sets for tiny arithmetic circuits or: How to turn bad designs into good

Quantum advantage and lower bounds in parallel query complexity

Counting Solutions to Conjunctive Queries: Structural and Hybrid Tractability

The tractability frontier of well-designed SPARQL queries

Querying Incomplete Data : Complexity and Tractability via Datalog and First-Order Rewritings

Direct sum theorems beyond query complexity

Tight Fine-Grained Bounds for Direct Access on Join Queries

Exponential Lower Bounds and Separation for Query Rewriting

Structure-Aware Lower Bounds and Broadening the Horizon of Tractability for QBF

Topology Dependent Bounds for FAQs.

Expected Shapley-Like Scores of Boolean Functions: Complexity and Applications to Probabilistic Databases

Preprocessing Complexity for Some Graph Problems Parameterized by Structural Parameters