Documentation

Unity Engine


User Manual

Script Reference

Unity Engine


X86.Avx

AVX intrinsics
Read time 26 minutesLast updated 4 days ago

Definition

public static class X86.Avx

Static Properties

Property

Description

IsAvxSupportedEvaluates to true at compile time if AVX intrinsics are supported.

Static Methods

Method

Description

broadcast_ssBroadcast a single-precision (32-bit) floating-point element from memory to all elements of dst.
cmp_pdCompare packed double-precision (64-bit) floating-point elements in a and b based on the comparison operand specified by imm8, and store the results in dst.
cmp_psCompare packed single-precision (32-bit) floating-point elements in a and b based on the comparison operand specified by imm8, and store the results in dst.
cmp_sdCompare the lower double-precision (64-bit) floating-point element in a and b based on the comparison operand specified by imm8, store the result in the lower element of dst, and copy the upper element from a to the upper element of dst.
cmp_ssCompare the lower single-precision (32-bit) floating-point element in a and b based on the comparison operand specified by imm8, store the result in the lower element of dst, and copy the upper 3 packed elements from a to the upper elements of dst.
maskload_pdLoad packed double-precision (64-bit) floating-point elements from memory into dst using mask (elements are zeroed out when the high bit of the corresponding element is not set).
maskload_psLoad packed single-precision (32-bit) floating-point elements from memory into dst using mask (elements are zeroed out when the high bit of the corresponding element is not set).
maskstore_pdStore packed double-precision (64-bit) floating-point elements from a into memory using mask.
maskstore_psStore packed single-precision (32-bit) floating-point elements from a into memory using mask.
mm256_add_pdAdd packed double-precision (64-bit) floating-point elements in a and b, and store the results in dst.
mm256_add_psAdd packed single-precision (32-bit) floating-point elements in a and b, and store the results in dst.
mm256_addsub_pdAlternatively add and subtract packed double-precision (64-bit) floating-point elements in a to/from packed elements in b, and store the results in dst.
mm256_addsub_psAlternatively add and subtract packed single-precision (32-bit) floating-point elements in a to/from packed elements in b, and store the results in dst.
mm256_and_pdCompute the bitwise AND of packed double-precision (64-bit) floating-point elements in a and b, and store the results in dst.
mm256_and_psCompute the bitwise AND of packed single-precision (32-bit) floating-point elements in a and b, and store the results in dst.
mm256_andnot_pdCompute the bitwise NOT of packed double-precision (64-bit) floating-point elements in a and then AND with b, and store the results in dst.
mm256_andnot_psCompute the bitwise NOT of packed single-precision (32-bit) floating-point elements in a and then AND with b, and store the results in dst.
mm256_blend_pdBlend packed double-precision (64-bit) floating-point elements from a and b using control mask imm8, and store the results in dst.
mm256_blend_psBlend packed single-precision (32-bit) floating-point elements from a and b using control mask imm8, and store the results in dst.
mm256_blendv_pdBlend packed double-precision (64-bit) floating-point elements from a and b using mask, and store the results in dst.
mm256_blendv_psBlend packed single-precision (32-bit) floating-point elements from a and b using mask, and store the results in dst.
mm256_broadcast_pdBroadcast 128 bits from memory (composed of 2 packed double-precision (64-bit) floating-point elements) to all elements of dst.
mm256_broadcast_psBroadcast 128 bits from memory (composed of 4 packed single-precision (32-bit) floating-point elements) to all elements of dst.
mm256_broadcast_sdBroadcast a double-precision (64-bit) floating-point element from memory to all elements of dst.
mm256_broadcast_ssBroadcast a single-precision (32-bit) floating-point element from memory to all elements of dst.
mm256_castpd128_pd256For compatibility with C++ code only. This is a no-op in Burst.
mm256_castpd256_pd128For compatibility with C++ code only. This is a no-op in Burst.
mm256_castpd_psFor compatibility with C++ code only. This is a no-op in Burst.
mm256_castpd_si256For compatibility with C++ code only. This is a no-op in Burst.
mm256_castps128_ps256For compatibility with C++ code only. This is a no-op in Burst.
mm256_castps256_ps128For compatibility with C++ code only. This is a no-op in Burst.
mm256_castps_pdFor compatibility with C++ code only. This is a no-op in Burst.
mm256_castps_si256For compatibility with C++ code only. This is a no-op in Burst.
mm256_castsi128_si256For compatibility with C++ code only. This is a no-op in Burst.
mm256_castsi256_pdFor compatibility with C++ code only. This is a no-op in Burst.
mm256_castsi256_psFor compatibility with C++ code only. This is a no-op in Burst.
mm256_castsi256_si128For compatibility with C++ code only. This is a no-op in Burst.
mm256_ceil_pdRound the packed double-precision (64-bit) floating-point elements in a up to an integer value, and store the results as packed double-precision floating-point elements in dst.
mm256_ceil_psRound the packed single-precision (32-bit) floating-point elements in a up to an integer value, and store the results as packed single-precision floating-point elements in dst.
mm256_cmp_pdCompare packed double-precision (64-bit) floating-point elements in a and b based on the comparison operand specified by imm8, and store the results in dst.
mm256_cmp_psCompare packed single-precision (32-bit) floating-point elements in a and b based on the comparison operand specified by imm8, and store the results in dst.
mm256_cvtepi32_pdConvert packed 32-bit integers in a to packed double-precision (64-bit) floating-point elements, and store the results in dst.
mm256_cvtepi32_psConvert packed 32-bit integers in a to packed single-precision (32-bit) floating-point elements, and store the results in dst.
mm256_cvtpd_epi32Convert packed double-precision(64-bit) floating-point elements in a to packed 32-bit integers, and store the results in dst.
mm256_cvtpd_psConvert packed double-precision (64-bit) floating-point elements in a to packed single-precision (32-bit) floating-point elements, and store the results in dst.
mm256_cvtps_epi32Convert packed single-precision (32-bit) floating-point elements in a to packed 32-bit integers, and store the results in dst.
mm256_cvtps_pdConvert packed single-precision (32-bit) floating-point elements in a to packed double-precision (64-bit) floating-point elements, and store the results in dst.
mm256_cvtss_f32Copy the lower single-precision (32-bit) floating-point element of a to dst.
mm256_cvttpd_epi32Convert packed double-precision (64-bit) floating-point elements in a to packed 32-bit integers with truncation, and store the results in dst.
mm256_cvttps_epi32Convert packed single-precision (32-bit) floating-point elements in a to packed 32-bit integers with truncation, and store the results in dst.
mm256_div_pdDivide packed double-precision (64-bit) floating-point elements in a by packed elements in b, and store the results in dst.
mm256_div_psDivide packed single-precision (32-bit) floating-point elements in a by packed elements in b, and store the results in dst.
mm256_dp_psConditionally multiply the packed single-precision (32-bit) floating-point elements in a and b using the high 4 bits in imm8, sum the four products, and conditionally store the sum in dst using the low 4 bits of imm8.
mm256_extract_epi32Extract a 32-bit integer from a, selected with index (which must be a constant), and store the result in dst.
mm256_extract_epi64Extract a 64-bit integer from a, selected with index (which must be a constant), and store the result in dst.
mm256_extractf128_pdExtract 128 bits (composed of 2 packed double-precision (64-bit) floating-point elements) from a, selected with imm8, and store the result in dst.
mm256_extractf128_psExtract 128 bits (composed of 4 packed single-precision (32-bit) floating-point elements) from a, selected with imm8, and store the result in dst.
mm256_extractf128_si256Extract 128 bits (composed of integer data) from a, selected with imm8, and store the result in dst.
mm256_floor_pdRound the packed double-precision (64-bit) floating-point elements in a down to an integer value, and store the results as packed double-precision floating-point elements in dst.
mm256_floor_psRound the packed single-precision (32-bit) floating-point elements in a down to an integer value, and store the results as packed single-precision floating-point elements in dst.
mm256_hadd_pdHorizontally add adjacent pairs of double-precision (64-bit) floating-point elements in a and b, and pack the results in dst.
mm256_hadd_psHorizontally add adjacent pairs of single-precision (32-bit) floating-point elements in a and b, and pack the results in dst.
mm256_hsub_pdHorizontally subtract adjacent pairs of double-precision (64-bit) floating-point elements in a and b, and pack the results in dst.
mm256_hsub_psHorizontally add adjacent pairs of single-precision (32-bit) floating-point elements in a and b, and pack the results in dst.
mm256_insert_epi16Copy a to dst, and insert the 16-bit integer i into dst at the location specified by index (which must be a constant).
mm256_insert_epi32Copy a to dst, and insert the 32-bit integer i into dst at the location specified by index (which must be a constant).
mm256_insert_epi64Copy a to dst, and insert the 64-bit integer i into dst at the location specified by index (which must be a constant).
mm256_insert_epi8Copy a to dst, and insert the 8-bit integer i into dst at the location specified by index (which must be a constant).
mm256_insertf128_pdCopy a to dst, then insert 128 bits (composed of 2 packed double-precision (64-bit) floating-point elements) from b into dst at the location specified by imm8.
mm256_insertf128_psCopy a to dst, then insert 128 bits (composed of 4 packed single-precision (32-bit) floating-point elements) from b into dst at the location specified by imm8.
mm256_insertf128_si256Copy a to dst, then insert 128 bits of integer data from b into dst at the location specified by imm8.
mm256_lddqu_si256Load 256-bits of integer data from unaligned memory into dst. This intrinsic may perform better than mm256_loadu_si256 when the data crosses a cache line boundary.
mm256_load_pdLoad 256-bits (composed of 8 packed single-precision (32-bit) floating-point elements) from memory
mm256_load_psLoad 256-bits (composed of 8 packed single-precision (32-bit) floating-point elements) from memory
mm256_load_si256Load 256-bits (composed of 8 packed 32-bit integers elements) from memory
mm256_loadu2_m128Load two 128-bit values (composed of 4 packed single-precision (32-bit) floating-point elements) from memory, and combine them into a 256-bit value in dst. hiaddr and loaddr do not need to be aligned on any particular boundary.
mm256_loadu2_m128dLoad two 128-bit values (composed of 2 packed double-precision (64-bit) floating-point elements) from memory, and combine them into a 256-bit value in dst. hiaddr and loaddr do not need to be aligned on any particular boundary.
mm256_loadu2_m128iLoad two 128-bit values (composed of integer data) from memory, and combine them into a 256-bit value in dst. hiaddr and loaddr do not need to be aligned on any particular boundary.
mm256_loadu_pdLoad 256-bits (composed of 4 packed double-precision (64-bit) floating-point elements) from memory
mm256_loadu_psLoad 256-bits (composed of 8 packed single-precision (32-bit) floating-point elements) from memory
mm256_loadu_si256Load 256-bits (composed of 8 packed 32-bit integers elements) from memory
mm256_maskload_pdLoad packed double-precision (64-bit) floating-point elements from memory into dst using mask (elements are zeroed out when the high bit of the corresponding element is not set).
mm256_maskload_psLoad packed single-precision (32-bit) floating-point elements from memory into dst using mask (elements are zeroed out when the high bit of the corresponding element is not set).
mm256_maskstore_pdStore packed double-precision (64-bit) floating-point elements from a into memory using mask.
mm256_maskstore_psStore packed single-precision (32-bit) floating-point elements from a into memory using mask.
mm256_max_pdCompare packed double-precision (64-bit) floating-point elements in a and b, and store packed maximum values in dst.
mm256_max_psCompare packed single-precision (32-bit) floating-point elements in a and b, and store packed maximum values in dst.
mm256_min_pdCompare packed double-precision (64-bit) floating-point elements in a and b, and store packed minimum values in dst.
mm256_min_psCompare packed single-precision (32-bit) floating-point elements in a and b, and store packed minimum values in dst.
mm256_movedup_pdDuplicate even-indexed double-precision (64-bit) floating-point elements from a, and store the results in dst.
mm256_movehdup_psDuplicate odd-indexed single-precision (32-bit) floating-point elements from a, and store the results in dst.
mm256_moveldup_psDuplicate even-indexed single-precision (32-bit) floating-point elements from a, and store the results in dst.
mm256_movemask_pdSet each bit of mask dst based on the most significant bit of the corresponding packed double-precision (64-bit) floating-point element in a.
mm256_movemask_psSet each bit of mask dst based on the most significant bit of the corresponding packed single-precision (32-bit) floating-point element in a.
mm256_mul_pdMultiply packed double-precision (64-bit) floating-point elements in a and b, and store the results in dst.
mm256_mul_psMultiply packed single-precision (32-bit) floating-point elements in a and b, and store the results in dst.
mm256_or_pdCompute the bitwise OR of packed double-precision (64-bit) floating-point elements in a and b, and store the results in dst.
mm256_or_psCompute the bitwise OR of packed single-precision (32-bit) floating-point elements in a and b, and store the results in dst.
mm256_permute2f128_pdShuffle 128-bits (composed of 2 packed double-precision (64-bit) floating-point elements) selected by imm8 from a and b, and store the results in dst.
mm256_permute2f128_psShuffle 128-bits (composed of 4 packed single-precision (32-bit) floating-point elements) selected by imm8 from a and b, and store the results in dst.
mm256_permute2f128_si256Shuffle 128-bits (composed of integer data) selected by imm8 from a and b, and store the results in dst.
mm256_permute_pdShuffle double-precision (64-bit) floating-point elements in a within 128-bit lanes using the control in imm8, and store the results in dst.
mm256_permute_psShuffle single-precision (32-bit) floating-point elements in a within 128-bit lanes using the control in imm8, and store the results in dst.
mm256_permutevar_pdShuffle double-precision (64-bit) floating-point elements in a within 128-bit lanes using the control in b, and store the results in dst.
mm256_permutevar_psShuffle single-precision (32-bit) floating-point elements in a within 128-bit lanes using the control in b, and store the results in dst.
mm256_rcp_psCompute the approximate reciprocal of packed single-precision (32-bit) floating-point elements in a, and store the results in dst. The maximum relative error for this approximation is less than 1.5*2^-12.
mm256_round_pdRound the packed double-precision (64-bit) floating-point elements in a using the rounding parameter, and store the results as packed double-precision floating-point elements in dst.
mm256_round_psRound the packed single-precision (32-bit) floating-point elements in a using the rounding parameter, and store the results as packed single-precision floating-point elements in dst.
mm256_rsqrt_psCompute the approximate reciprocal square root of packed single-precision (32-bit) floating-point elements in a, and store the results in dst. The maximum relative error for this approximation is less than 1.5*2^-12.
mm256_set1_epi16Broadcast 16-bit integer a to all all elements of dst. This intrinsic may generate the vpbroadcastw instruction.
mm256_set1_epi32Broadcast 32-bit integer a to all elements of dst. This intrinsic may generate the vpbroadcastd instruction.
mm256_set1_epi64xBroadcast 64-bit integer a to all elements of dst. This intrinsic may generate the vpbroadcastq instruction.
mm256_set1_epi8Broadcast 8-bit integer a to all elements of dst. This intrinsic may generate the vpbroadcastb instruction.
mm256_set1_pdBroadcast double-precision (64-bit) floating-point value a to all elements of dst.
mm256_set1_psBroadcast single-precision (32-bit) floating-point value a to all elements of dst.
mm256_set_epi16Set packed short elements in dst with the supplied values.
mm256_set_epi32Set packed int elements in dst with the supplied values.
mm256_set_epi64xSet packed 64-bit integers in dst with the supplied values.
mm256_set_epi8Set packed byte elements in dst with the supplied values.
mm256_set_m128Set packed __m256 vector dst with the supplied values.
mm256_set_m128dSet packed v256 vector with the supplied values.
mm256_set_m128iSet packed v256 vector with the supplied values.
mm256_set_pdSet packed double-precision (64-bit) floating-point elements in dst with the supplied values.
mm256_set_psSet packed single-precision (32-bit) floating-point elements in dst with the supplied values.
mm256_setr_epi16Set packed short elements in dst with the supplied values in reverse order.
mm256_setr_epi32Set packed int elements in dst with the supplied values in reverse order.
mm256_setr_epi64xSet packed 64-bit integers in dst with the supplied values in reverse order.
mm256_setr_epi8Set packed byte elements in dst with the supplied values in reverse order.
mm256_setr_m128Set packed v256 vector with the supplied values in reverse order.
mm256_setr_m128dSet packed v256 vector with the supplied values in reverse order.
mm256_setr_m128iSet packed v256 vector with the supplied values in reverse order.
mm256_setr_pdSet packed double-precision (64-bit) floating-point elements in dst with the supplied values in reverse order.
mm256_setr_psSet packed single-precision (32-bit) floating-point elements in dst with the supplied values in reverse order.
mm256_setzero_pdReturn Vector with all elements set to zero.
mm256_setzero_psReturn Vector with all elements set to zero.
mm256_setzero_si256Return Vector with all elements set to zero.
mm256_shuffle_pdShuffle double-precision (64-bit) floating-point elements within 128-bit lanes using the control in imm8, and store the results in dst.
mm256_shuffle_psShuffle single-precision (32-bit) floating-point elements in a within 128-bit lanes using the control in imm8, and store the results in dst.
mm256_sqrt_pdCompute the square root of packed double-precision (64-bit) floating-point elements in a, and store the results in dst.
mm256_sqrt_psCompute the square root of packed single-precision (32-bit) floating-point elements in a, and store the results in dst.
mm256_store_pdStore 256-bits (composed of 4 packed double-precision (64-bit) floating-point elements) from a into memory
mm256_store_psStore 256-bits (composed of 8 packed single-precision (32-bit) floating-point elements) from a into memory
mm256_store_si256Store 256-bits (composed of 8 packed 32-bit integer elements) from a into memory
mm256_storeu2_m128Store the high and low 128-bit halves (each composed of 4 packed single-precision (32-bit) floating-point elements) from a into memory two different 128-bit locations. hiaddr and loaddr do not need to be aligned on any particular boundary.
mm256_storeu2_m128dStore the high and low 128-bit halves (each composed of 2 packed double-precision (64-bit) floating-point elements) from a into memory two different 128-bit locations. hiaddr and loaddr do not need to be aligned on any particular boundary.
mm256_storeu2_m128iStore the high and low 128-bit halves (each composed of integer data) from a into memory two different 128-bit locations. hiaddr and loaddr do not need to be aligned on any particular boundary.
mm256_storeu_pdStore 256-bits (composed of 4 packed double-precision (64-bit) floating-point elements) from a into memory
mm256_storeu_psStore 256-bits (composed of 8 packed single-precision (32-bit) floating-point elements) from a into memory
mm256_storeu_si256Store 256-bits (composed of 8 packed 32-bit integer elements) from a into memory
mm256_stream_pdStore 256-bits (composed of 4 packed double-precision (64-bit) floating-point elements) from a into memory using a non-temporal memory hint. mem_addr must be aligned on a 32-byte boundary or a general-protection exception may be generated.
mm256_stream_psStore 256-bits (composed of 8 packed single-precision (32-bit) floating-point elements) from a into memory using a non-temporal memory hint. mem_addr must be aligned on a 32-byte boundary or a general-protection exception may be generated.
mm256_stream_si256Store 256-bits of integer data from a into memory using a non-temporal memory hint. mem_addr must be aligned on a 32-byte boundary or a general-protection exception may be generated.
mm256_sub_pdSubtract packed double-precision (64-bit) floating-point elements in b from packed double-precision (64-bit) floating-point elements in a, and store the results in dst.
mm256_sub_psSubtract packed single-precision (32-bit) floating-point elements in b from packed single-precision (32-bit) floating-point elements in a, and store the results in dst.
mm256_testc_pdCompute the bitwise AND of 256 bits (representing double-precision (64-bit) floating-point elements) in a and b, producing an intermediate 256-bit value, and set ZF to 1 if the sign bit of each 64-bit element in the intermediate value is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, producing an intermediate value, and set CF to 1 if the sign bit of each 64-bit element in the intermediate value is zero, otherwise set CF to 0. Return the CF value.
mm256_testc_psCompute the bitwise AND of 256 bits (representing single-precision (32-bit) floating-point elements) in a and b, producing an intermediate 256-bit value, and set ZF to 1 if the sign bit of each 32-bit element in the intermediate value is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, producing an intermediate value, and set CF to 1 if the sign bit of each 32-bit element in the intermediate value is zero, otherwise set CF to 0. Return the CF value.
mm256_testc_si256Compute the bitwise AND of 256 bits (representing integer data) in a and b, and set ZF to 1 if the result is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, and set CF to 1 if the result is zero, otherwise set CF to 0. Return the CF value.
mm256_testnzc_pdCompute the bitwise AND of 256 bits (representing double-precision (64-bit) floating-point elements) in a and b, producing an intermediate 256-bit value, and set ZF to 1 if the sign bit of each 64-bit element in the intermediate value is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, producing an intermediate value, and set CF to 1 if the sign bit of each 64-bit element in the intermediate value is zero, otherwise set CF to 0. Return 1 if both the ZF and CF values are zero, otherwise return 0.
mm256_testnzc_psCompute the bitwise AND of 256 bits (representing single-precision (32-bit) floating-point elements) in a and b, producing an intermediate 256-bit value, and set ZF to 1 if the sign bit of each 32-bit element in the intermediate value is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, producing an intermediate value, and set CF to 1 if the sign bit of each 32-bit element in the intermediate value is zero, otherwise set CF to 0. Return 1 if both the ZF and CF values are zero, otherwise return 0.
mm256_testnzc_si256Compute the bitwise AND of 256 bits (representing integer data) in a and b, and set ZF to 1 if the result is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, and set CF to 1 if the result is zero, otherwise set CF to 0. Return 1 if both the ZF and CF values are zero, otherwise return 0.
mm256_testz_pdCompute the bitwise AND of 256 bits (representing double-precision (64-bit) floating-point elements) in a and b, producing an intermediate 256-bit value, and set ZF to 1 if the sign bit of each 64-bit element in the intermediate value is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, producing an intermediate value, and set CF to 1 if the sign bit of each 64-bit element in the intermediate value is zero, otherwise set CF to 0. Return the ZF value.
mm256_testz_psCompute the bitwise AND of 256 bits (representing single-precision (32-bit) floating-point elements) in a and b, producing an intermediate 256-bit value, and set ZF to 1 if the sign bit of each 32-bit element in the intermediate value is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, producing an intermediate value, and set CF to 1 if the sign bit of each 32-bit element in the intermediate value is zero, otherwise set CF to 0. Return the ZF value.
mm256_testz_si256Compute the bitwise AND of 256 bits (representing integer data) in a and b, and set ZF to 1 if the result is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, and set CF to 1 if the result is zero, otherwise set CF to 0. Return the ZF value.
mm256_undefined_pdReturn a 256-bit vector with undefined contents.
mm256_undefined_psReturn a 256-bit vector with undefined contents.
mm256_undefined_si256Return a 256-bit vector with undefined contents.
mm256_unpackhi_pdUnpack and interleave double-precision (64-bit) floating-point elements from the high half of each 128-bit lane in a and b, and store the results in dst.
mm256_unpackhi_psUnpack and interleave single-precision(32-bit) floating-point elements from the high half of each 128-bit lane in a and b, and store the results in dst.
mm256_unpacklo_pdUnpack and interleave double-precision (64-bit) floating-point elements from the low half of each 128-bit lane in a and b, and store the results in dst.
mm256_unpacklo_psUnpack and interleave single-precision (32-bit) floating-point elements from the low half of each 128-bit lane in a and b, and store the results in dst.
mm256_xor_pdCompute the bitwise XOR of packed double-precision (64-bit) floating-point elements in a and b, and store the results in dst.
mm256_xor_psCompute the bitwise XOR of packed single-precision (32-bit) floating-point elements in a and b, and store the results in dst.
mm256_zeroallZeros the contents of all YMM registers
mm256_zeroupperZero the upper 128 bits of all YMM registers; the lower 128-bits of the registers are unmodified.
mm256_zextpd128_pd256Casts vector of type v128 to type v256; the upper 128 bits of the result are zeroed. This intrinsic is only used for compilation and does not generate any instructions, thus it has zero latency.
mm256_zextps128_ps256Casts vector of type v128 to type v256; the upper 128 bits of the result are zeroed. This intrinsic is only used for compilation and does not generate any instructions, thus it has zero latency.
mm256_zextsi128_si256Casts vector of type v128 to type v256; the upper 128 bits of the result are zeroed. This intrinsic is only used for compilation and does not generate any instructions, thus it has zero latency.
permute_pdShuffle double-precision (64-bit) floating-point elements in a using the control in imm8, and store the results in dst.
permute_psShuffle single-precision (32-bit) floating-point elements in a using the control in imm8, and store the results in dst.
permutevar_pdShuffle double-precision (64-bit) floating-point elements in a using the control in b, and store the results in dst.
permutevar_psShuffle single-precision (32-bit) floating-point elements in a using the control in b, and store the results in dst.
testc_pdCompute the bitwise AND of 128 bits (representing double-precision (64-bit) floating-point elements) in a and b, producing an intermediate 128-bit value, and set ZF to 1 if the sign bit of each 64-bit element in the intermediate value is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, producing an intermediate value, and set CF to 1 if the sign bit of each 64-bit element in the intermediate value is zero, otherwise set CF to 0. Return the CF value.
testc_psCompute the bitwise AND of 128 bits (representing single-precision (32-bit) floating-point elements) in a and b, producing an intermediate 128-bit value, and set ZF to 1 if the sign bit of each 32-bit element in the intermediate value is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, producing an intermediate value, and set CF to 1 if the sign bit of each 32-bit element in the intermediate value is zero, otherwise set CF to 0. Return the CF value.
testnzc_pdCompute the bitwise AND of 128 bits (representing double-precision (64-bit) floating-point elements) in a and b, producing an intermediate 128-bit value, and set ZF to 1 if the sign bit of each 64-bit element in the intermediate value is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, producing an intermediate value, and set CF to 1 if the sign bit of each 64-bit element in the intermediate value is zero, otherwise set CF to 0. Return 1 if both the ZF and CF values are zero, otherwise return 0.
testnzc_psCompute the bitwise AND of 128 bits (representing single-precision (32-bit) floating-point elements) in a and b, producing an intermediate 128-bit value, and set ZF to 1 if the sign bit of each 32-bit element in the intermediate value is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, producing an intermediate value, and set CF to 1 if the sign bit of each 32-bit element in the intermediate value is zero, otherwise set CF to 0. Return 1 if both the ZF and CF values are zero, otherwise return 0.
testz_pdCompute the bitwise AND of 128 bits (representing double-precision (64-bit) floating-point elements) in a and b, producing an intermediate 128-bit value, and set ZF to 1 if the sign bit of each 64-bit element in the intermediate value is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, producing an intermediate value, and set CF to 1 if the sign bit of each 64-bit element in the intermediate value is zero, otherwise set CF to 0. Return the ZF value.
testz_psCompute the bitwise AND of 128 bits (representing single-precision (32-bit) floating-point elements) in a and b, producing an intermediate 128-bit value, and set ZF to 1 if the sign bit of each 32-bit element in the intermediate value is zero, otherwise set ZF to 0. Compute the bitwise NOT of a and then AND with b, producing an intermediate value, and set CF to 1 if the sign bit of each 32-bit element in the intermediate value is zero, otherwise set CF to 0. Return the ZF value.
undefined_pdReturn a 128-bit vector with undefined contents.
undefined_psReturn a 128-bit vector with undefined contents.
undefined_si128Return a 128-bit vector with undefined contents.

Enums

Enum

Description

X86.Avx.CMPCompare predicates for scalar and packed compare intrinsic functions