eigen - C++ library for linear algebra

	Commit message (Collapse)	Author	Age
*	Fix ambiguity due to argument dependent lookup.	Nathan Luehr	2021-05-11
\|
*	Fix pow and other cwise ops for half/bfloat16.	Antonio Sanchez	2021-01-22
\| \| \| \| \| \| \| \| \| \| \| \| \|	The new `generic_pow` implementation was failing for half/bfloat16 since their construction from int/float is not `constexpr`. Modified in `GenericPacketMathFunctions` to remove `constexpr`. While adding tests for half/bfloat16, found other issues related to implicit conversions. Also needed to implement `numext::arg` for non-integer, non-complex, non-float/double/long double types. These seem to be implicitly converted to `std::complex<T>`, which then fails for half/bfloat16.
*	Improved std::complex sqrt and rsqrt.	Antonio Sanchez	2021-01-17
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Replaces `std::sqrt` with `complex_sqrt` for all platforms (previously `complex_sqrt` was only used for CUDA and MSVC), and implements custom `complex_rsqrt`. Also introduces `numext::rsqrt` to simplify implementation, and modified `numext::hypot` to adhere to IEEE IEC 6059 for special cases. The `complex_sqrt` and `complex_rsqrt` implementations were found to be significantly faster than `std::sqrt<std::complex<T>>` and `1/numext::sqrt<std::complex<T>>`. Benchmark file attached. ``` GCC 10, Intel Xeon, x86_64: --------------------------------------------------------------------------- Benchmark Time CPU Iterations --------------------------------------------------------------------------- BM_Sqrt<std::complex<float>> 9.21 ns 9.21 ns 73225448 BM_StdSqrt<std::complex<float>> 17.1 ns 17.1 ns 40966545 BM_Sqrt<std::complex<double>> 8.53 ns 8.53 ns 81111062 BM_StdSqrt<std::complex<double>> 21.5 ns 21.5 ns 32757248 BM_Rsqrt<std::complex<float>> 10.3 ns 10.3 ns 68047474 BM_DivSqrt<std::complex<float>> 16.3 ns 16.3 ns 42770127 BM_Rsqrt<std::complex<double>> 11.3 ns 11.3 ns 61322028 BM_DivSqrt<std::complex<double>> 16.5 ns 16.5 ns 42200711 Clang 11, Intel Xeon, x86_64: --------------------------------------------------------------------------- Benchmark Time CPU Iterations --------------------------------------------------------------------------- BM_Sqrt<std::complex<float>> 7.46 ns 7.45 ns 90742042 BM_StdSqrt<std::complex<float>> 16.6 ns 16.6 ns 42369878 BM_Sqrt<std::complex<double>> 8.49 ns 8.49 ns 81629030 BM_StdSqrt<std::complex<double>> 21.8 ns 21.7 ns 31809588 BM_Rsqrt<std::complex<float>> 8.39 ns 8.39 ns 82933666 BM_DivSqrt<std::complex<float>> 14.4 ns 14.4 ns 48638676 BM_Rsqrt<std::complex<double>> 9.83 ns 9.82 ns 70068956 BM_DivSqrt<std::complex<double>> 15.7 ns 15.7 ns 44487798 Clang 9, Pixel 2, aarch64: --------------------------------------------------------------------------- Benchmark Time CPU Iterations --------------------------------------------------------------------------- BM_Sqrt<std::complex<float>> 24.2 ns 24.1 ns 28616031 BM_StdSqrt<std::complex<float>> 104 ns 103 ns 6826926 BM_Sqrt<std::complex<double>> 31.8 ns 31.8 ns 22157591 BM_StdSqrt<std::complex<double>> 128 ns 128 ns 5437375 BM_Rsqrt<std::complex<float>> 31.9 ns 31.8 ns 22384383 BM_DivSqrt<std::complex<float>> 99.2 ns 98.9 ns 7250438 BM_Rsqrt<std::complex<double>> 46.0 ns 45.8 ns 15338689 BM_DivSqrt<std::complex<double>> 119 ns 119 ns 5898944 ```
*	Replace M_LOG2E and M_LN2 with custom macros.	Antonio Sanchez	2020-12-11
\| \| \| \| \| \| \| \| \| \|	For these to exist we would need to define `_USE_MATH_DEFINES` before `cmath` or `math.h` is first included. However, we don't control the include order for projects outside Eigen, so even defining the macro in `Eigen/Core` does not fix the issue for projects that end up including `<cmath>` before Eigen does (explicitly or transitively). To fix this, we define `EIGEN_LOG2E` and `EIGEN_LN2` ourselves.
*	Add log2() to Eigen.	Rasmus Munk Larsen	2020-12-04
\|
*	Revert "Add log2() operator to Eigen"	Rasmus Munk Larsen	2020-12-03
\| \| \| \|	This reverts commit 4d91519a9be061da5d300079fca17dd0b9328050.
*	Add log2() operator to Eigen	Rasmus Munk Larsen	2020-12-03
\|
*	Fix boolean float conversion and product warnings.	Antonio Sanchez	2020-11-24
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This fixes some gcc warnings such as: ``` Eigen/src/Core/GenericPacketMath.h:655:63: warning: implicit conversion turns floating-point number into bool: 'typename __gnu_cxx::__enable_if<__is_integer<bool>::__value, double>::__type' (aka 'double') to 'bool' [-Wimplicit-conversion-floating-point-to-bool] Packet psqrt(const Packet& a) { EIGEN_USING_STD(sqrt); return sqrt(a); } ``` Details: - Added `scalar_sqrt_op<bool>` (`-Wimplicit-conversion-floating-point-to-bool`). - Added `scalar_square_op<bool>` and `scalar_cube_op<bool>` specializations (`-Wint-in-bool-context`) - Deprecated above specialized ops for bool. - Modified `cxx11_tensor_block_eval` to specialize generator for booleans (`-Wint-in-bool-context`) and to use `abs` instead of `square` to avoid deprecated bool ops.
*	Drop EIGEN_USING_STD_MATH in favour of EIGEN_USING_STD	David Tellenbach	2020-10-09
\|
*	Change the sign operator in Eigen to return NaN for NaN arguments, not zero.	Rasmus Munk Larsen	2020-07-07
\|
*	Fix compilation error in logistic packet op.	Rasmus Munk Larsen	2020-06-03
\|
*	Bug #1777: make the scalar and packet path consistent for the logistic ↵	Gael Guennebaud	2020-05-31
\| \| \| \|	function + respective unit test
*	Remove reference to non-existent unary_op_base class.	Rasmus Munk Larsen	2020-03-19
\|
*	Add shift_left<N> and shift_right<N> coefficient-wise unary Array functions	Joel Holdsworth	2020-03-19
\|
*	Don't use the rational approximation to the logistic function on GPUs as it ↵	Rasmus Munk Larsen	2020-01-09
\| \| \| \|	appears to be slightly slower.
*	The upper limits for where to use the rational approximation to the logistic ↵	Rasmus Munk Larsen	2020-01-08
\| \| \| \|	function were not set carefully enough in the original commit, and some arguments would cause the function to return values greater than 1. This change set the versions found by scanning all floating point numbers (using std::nextafterf()).
*	Bug #1785: Introduce numext::rint.	Ilya Tokar	2020-01-07
\| \| \| \| \| \|	This provides a new op that matches std::rint and previous behavior of pround. Also adds corresponding unsupported/../Tensor op. Performance is the same as e. g. floor (tested SSE/AVX).
*	Improve accuracy of fast approximate tanh and the logistic functions in ↵	Rasmus Munk Larsen	2019-12-16
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Eigen, such that they preserve relative accuracy to within a few ULPs where their function values tend to zero (around x=0 for tanh, and for large negative x for the logistic function). This change re-instates the fast rational approximation of the logistic function for float32 in Eigen (removed in https://gitlab.com/libeigen/eigen/commit/66f07efeaed39d6a67005343d7e0caf7d9eeacdb), but uses the more accurate approximation 1/(1+exp(-1)) ~= exp(x) below -9. The exponential is only calculated on the vectorized path if at least one element in the SIMD input vector is less than -9. This change also contains a few improvements to speed up the original float specialization of logistic: - Introduce EIGEN_PREDICT_{FALSE,TRUE} for __builtin_predict and use it to predict that the logistic-only path is most likely (~2-3% speedup for the common case). - Carefully set the upper clipping point to the smallest x where the approximation evaluates to exactly 1. This saves the explicit clamping of the output (~7% speedup). The increased accuracy for tanh comes at a cost of 10-20% depending on instruction set. The benchmarks below repeated calls u = v.logistic() (u = v.tanh(), respectively) where u and v are of type Eigen::ArrayXf, have length 8k, and v contains random numbers in [-1,1]. Benchmark numbers for logistic: Before: Benchmark Time(ns) CPU(ns) Iterations ----------------------------------------------------------------- SSE BM_eigen_logistic_float 4467 4468 155835 model_time: 4827 AVX BM_eigen_logistic_float 2347 2347 299135 model_time: 2926 AVX+FMA BM_eigen_logistic_float 1467 1467 476143 model_time: 2926 AVX512 BM_eigen_logistic_float 805 805 858696 model_time: 1463 After: Benchmark Time(ns) CPU(ns) Iterations ----------------------------------------------------------------- SSE BM_eigen_logistic_float 2589 2590 270264 model_time: 4827 AVX BM_eigen_logistic_float 1428 1428 489265 model_time: 2926 AVX+FMA BM_eigen_logistic_float 1059 1059 662255 model_time: 2926 AVX512 BM_eigen_logistic_float 673 673 1000000 model_time: 1463 Benchmark numbers for tanh: Before: Benchmark Time(ns) CPU(ns) Iterations ----------------------------------------------------------------- SSE BM_eigen_tanh_float 2391 2391 292624 model_time: 4242 AVX BM_eigen_tanh_float 1256 1256 554662 model_time: 2633 AVX+FMA BM_eigen_tanh_float 823 823 866267 model_time: 1609 AVX512 BM_eigen_tanh_float 443 443 1578999 model_time: 805 After: Benchmark Time(ns) CPU(ns) Iterations ----------------------------------------------------------------- SSE BM_eigen_tanh_float 2588 2588 273531 model_time: 4242 AVX BM_eigen_tanh_float 1536 1536 452321 model_time: 2633 AVX+FMA BM_eigen_tanh_float 1007 1007 694681 model_time: 1609 AVX512 BM_eigen_tanh_float 471 471 1472178 model_time: 805
*	Revert the specialization for scalar_logistic_op<float> introduced in:	Rasmus Munk Larsen	2019-12-02
\| \| \| \| \| \| \|	https://bitbucket.org/eigen/eigen/commits/77b447c24e3344e43ff64eb932d4bb35a2db01ce While providing a 50% speedup on Haswell+ processors, the large relative error outside [-18, 18] in this approximation causes problems, e.g., when computing gradients of activation functions like softplus in neural networks.
*	PR 719: fix real/imag namespace conflict	Gael Guennebaud	2019-10-08
\|
*	[SYCL] This PR adds the minimum modifications to Eigen core required to run ↵	Mehdi Goli	2019-06-27
\| \| \| \| \| \| \| \|	Eigen unsupported modules on devices supporting SYCL. * Adding SYCL memory model * Enabling/Disabling SYCL backend in Core * Supporting Vectorization
*	Fix build with clang on Windows.	Rasmus Munk Larsen	2019-05-09
\|
*	Restore C++03 compatibility	Christoph Hertzberg	2019-05-06
\|
*	Fix traits for scalar_logistic_op.	Rasmus Munk Larsen	2019-05-03
\|
*	Make clipping outside [-18:18] consistent for vectorized and non-vectorized ↵	Rasmus Munk Larsen	2019-03-15
\| \| \| \|	paths of scalar_logistic_<float>.
*	Set cost of conjugate to 0 (in practice it boils down to a no-op).	Gael Guennebaud	2019-02-18
\| \| \| \| \|	This is also important to make sure that A.conjugate() * B.conjugate() does not evaluate its arguments into temporaries (e.g., if A and B are fixed and small, or * fall back to lazyProduct)
*	Add support for inverse hyperbolic functions.	Rasmus Munk Larsen	2019-01-11
\| \| \| \|	Fix cost of division.
*	Add optimized version of logistic function for float. As an example, this is ↵	Rasmus Munk Larsen	2018-11-12
\| \| \| \|	about 50% faster than the existing version on Haswell using AVX.
*	sigmoid -> logistic	Rasmus Munk Larsen	2018-08-13
\|
*	Move sigmoid functor to core.	Rasmus Munk Larsen	2018-08-03
\|
*	Added support for expm1 in Eigen.	Srinivas Vasudevan	2016-12-02
\|
*	Added isnan, isfinite and isinf for SYCL device. Plus test for that.	Luke Iwanski	2016-11-18
\|
*	Adding EIGEN_DEVICE_FUNC in the Geometry module.	Robert Lukierski	2016-10-12
\| \| \| \| \|	Additional CUDA necessary fixes in the Core (mostly usage of EIGEN_USING_STD_MATH).
*	bug #1195: move NumTraits::Div<>::Cost to internal::scalar_div_cost (with ↵	Gael Guennebaud	2016-09-08
\| \| \| \|	some specializations in arch/SSE and arch/AVX)
*	Cleanup cost of tanh	Gael Guennebaud	2016-08-23
\|
*	Factorize the 4 copies of tanh implementations, make numext::tanh consistent ↵	Gael Guennebaud	2016-08-23
\| \| \| \|	with array::tanh, enable fast tanh in fast-math mode only.
*	fix tanh inconsistent	Ziming Dong	2016-08-06
\|
*	bug #1232: refactor special functions as a new SpecialFunctions module, ↵	Gael Guennebaud	2016-07-08
\| \| \| \|	currently in unsupported/.
*	Expose log1p to Array.	Gael Guennebaud	2016-06-01
\|
*	Roll back changes to core. Move include of TensorFunctors.h up to satisfy ↵	Rasmus Munk Larsen	2016-05-17
\| \| \| \|	dependence in TensorCostModel.h.
*	Improvements to parallelFor.	Rasmus Munk Larsen	2016-05-12
\| \| \| \|	Move some scalar functors from TensorFunctors. to Eigen core.
*	Fixed a typo in my previous commit	Benoit Steiner	2016-05-11
\|
*	Enable and fix -Wdouble-conversion warnings	Christoph Hertzberg	2016-05-05
\|
*	Improved support for trigonometric functions on GPU	Benoit Steiner	2016-04-13
\|
*	Added support for computing cos, sin, tan, and tanh on GPU.	Benoit Steiner	2016-04-13
\|
*	Don't put a command at the end of an enumerator list	Benoit Steiner	2016-04-12
\|
*	More accurate cost estimates for exp, log, tanh, and sqrt.	Benoit Steiner	2016-04-11
\|
*	Updated the unary functors to use the numext implementation of typicall ↵	Benoit Steiner	2016-04-07
\| \| \| \|	functions instead of the one provided in the standard library. The standard library functions aren't supported officially by cuda, so we're better off using the numext implementations.
*	Merged eigen/eigen into default	tillahoffmann	2016-04-05
\|\
\| *	Updated the scalar_abs_op struct to make it compatible with cuda devices.	Benoit Steiner	2016-04-04
\| \|