Is there a trick to cut down the overhead when amplitude-encoding a batch of real-valued vectors?
I tried a few different normalization schemes on a 4-qubit test case and the preparation unitary still dominated the total gate count before the variational part even ran. Not sure whether a block-encoding approach would actually move the bottleneck or just shift it elsewhere.