by Alex Mulrooney, Zhi Li, Austin J. BrockmeierPredicting the neural response to natural images in the visual cortex requires extracting relevant features from the images and relating those feature to the observed responses. In this work, we optimize the feature extraction in order to maximize the information shared between the image features and the neural response across voxels in a given region of interest (ROI) extracted from the BOLD signal measured by functional magnetic resonance imaging (fMRI). We adapt contrastive learning (CL) to fine-tune a convolutional neural network, which was pretrained for image classification, such that a mapping of a given image’s features are more similar to the corresponding fMRI response than to the responses to other images. We exploit the Natural Scenes Dataset as organized for the Algonauts Project, which contains the high-resolution fMRI responses of eight subjects to tens of thousands of naturalistic images. We show that CL fine-tuning creates feature extraction models that enable higher encoding accuracy in both early and higher visual ROIs as compared to the features from the pretrained network. Quantitatively, performance is similar to baseline approach that directly uses a regression loss at the output of the network to tune it for fMRI response encoding. We investigate inter-subject transfer of the CL fine-tuned models, including subjects from the Natural Object Dataset, another lower-resolution dataset with 9 subjects. We also pool subjects for fine-tuning, which further improves encoding performance in early ROIs. Finally, we examine the performance of the fine-tuned models on common image classification tasks, explore the landscape of ROI-specific models by applying dimensionality reduction on the Bhattacharya dissimilarity matrix created using the predictions on those tasks, show that these landscapes match those based on representational similarity analysis. Finally, we generate images via Stable Diffusion based on vector-space prompts created by aligning the CL-tuned models embeddings for different ROIs; showing that generated images have similar embeddings to the original but that estimates of the intrinsic dimension are lower for generated versus original representations.