Have you tried this?
// Modular reduction: (uint64(m) * uint64(max)) >> 32
// first multiply (even lanes)
VPMULUDQ Z31, Z2, Z3
// prepare odd lanes multiply
VPSRLQ $32, Z3, Z3
VPSRLQ $32, Z2, Z2
// second multiply (odd lanes)
VPMULUDQ Z31, Z2, Z2
// clear wrong lane
VPSRLQ $32, Z2, Z2
VPSLLQ $32, Z2, Z2
// combine odd and even lanes
VPORQ Z2, Z3, Z3
// Store result
VMOVDQU32 Z3, (DI)