Existing supervised and self-supervised EEG models mainly learn discriminative or reconstructive representations within individual segments, while the transition information between adjacent EEG segments remains underexplored. In this study, we propose a Multimodal self-supervised EEG World Model for wearable seizure detection. Inspired by Le World Model, the proposed method encodes consecutive EEG segments into a shared latent space and predicts the next-segment latent representation from the current-segment representation conditioned on synchronized physiological information from ECG, EMG, and movement (MOV) signals. A learnable query-based fusion module aggregates the auxiliary multimodal representations into a compact physiological condition, while Sketched Isotropic Gaussian Regularization (SIGReg) is applied to stabilize the latent space and prevent representation collapse. After pretraining, only the pretrained EEG encoder is retained and frozen for linear binary probing, enabling EEG-only downstream seizure detection. We evaluated the proposed model on the SeizeIT2 wearable focal epilepsy dataset using a strict patient-wise training, validation, and test split. The proposed Multimodal EEG World Model achieved an AUPRC of 0.3748 , ROC-AUC of 0.8025 , and balanced accuracy of 0.7308 , ranking first on these three metrics among the ablation studies. It also achieved the highest AUPRC, ROC-AUC, balanced accuracy, and F1-score among the evaluated external baselines. These findings demonstrate that synchronized multimodal physiological information can provide useful contextual information for latent EEG transition learning and improve wearable EEG representation learning.