Fairness audits of chest radiograph artificial intelligence usually score subgroup performance against labels extracted from radiology reports by natural language processing (NLP). We tested whether these labels can make a model look fairer than it is, a failure we call phantom fairness. On 4,376 NIH ChestX-ray14 images with both NLP and radiologist-adjudicated labels, we audited classifiers from three backbones by sex, age, projection and older women. NLP labels missed 54% to 59% of radiologist-confirmed airspace opacity, pneumothorax and nodule or mass, and every fracture. These misses were not random: models scored them lower than NLP-caught positives for every finding and backbone (median rank-biserial correlation 0.20 for the primary backbone and 0.16 for the others; all adjusted p [≤] 0.038), so the labelling errors hid the models' own mistakes. Subgroup gaps were larger on adjudicated labels in 16 of 24 comparisons, none individually significant. For pneumothorax in women older than 60 years, the NLP audit showed a sensitivity gap of 0.064 against a true gap of 0.184; removing the hardest positives reproduced this concealment (0.124 against 0.120 observed), whereas random label errors produced less (0.077). Fairness audits based on NLP-extracted labels should be treated as provisional until checked against radiologist-adjudicated labels.