1. Why use the concatenated multi-layer features from the visual encoder as the KV for the text query? 2. Do the selected specific layer visual features have any special effects or basis? Thank you!
Thank you!