Breast image classification requires local detail and global tissue context, yet these cues can weaken as representations deepen.
We present M3D-Net, a mammography encoder that hierarchically coordinates multi-scale coordinate attention, bounded dynamic feature reuse, and differential attention through resolution-aware operator placement.
Within-stage retrieval preserves access to earlier features, coordinate-aware aggregation integrates local and global context, and differential attention operates at coarse resolutions.
We evaluate image-only classification on AISSLab mammography and an adapted image--clinical model on BrEaST ultrasound.
Against EdgeNeXt, RepViT, and TransXNet, the proposed implementations achieve the highest recorded validation accuracy and late-training accuracy, with the lowest endpoint cross-entropy loss.
Validation accuracies reach 97.78% and 80.39%, respectively.
These results support further evaluation of hierarchical coordination across breast imaging settings; repeated-seed, component-controlled, and independent evaluations remain necessary.