7eb0945af254d376df11475150d184623104cf93 - platform/external/skia

commit	7eb0945af254d376df11475150d184623104cf93	[log] [tgz]
author	mtklein <mtklein@chromium.org>	Fri Jul 31 10:46:50 2015 -0700
committer	Commit bot <commit-bot@chromium.org>	Fri Jul 31 10:46:50 2015 -0700
tree	f1eca7b5e83428fc5f728967fed41324d5186e66
parent	5119ac069e6cf70175b5581eedee7d07347b216a [diff]

Port SkUtils opts to SkOpts.

With this new arrangement, the benefits of inlining sk_memset16/32 have changed.

On x86, they're not significantly different, except for small N<=10 where the inlined code is significantly slower.
On ARMv7 with NEON, our custom code is still significantly faster for N>10 (up to 2x faster).  For small N<=10 inlining is still significantly faster.
On ARMv7 without NEON, our custom code is still ridiculously faster (up to 10x) than inlining for N>10, though for small N<=10 inlining is still a little faster.

We were not using the NEON memset16 and memset32 procs on ARMv8.  At first blush, that seems to be an oversight, but if so it's an extremely lucky one.  The ARMv8 code generation for our memset16/32 procs is total garbage, leaving those methods ~8x slower than just inlining the memset, using the compiler's autovectorization.

So, no need to inline any more on x86, and still inline for N<=10 on ARMv7.  Always inline for ARMv8.

BUG=skia:4117

Review URL: https://codereview.chromium.org/1270573002

gyp/opts.gypi[diff]
include/core/SkFloatingPoint.h[diff]
include/core/SkUtils.h[diff]
include/private/SkOpts.h[Renamed from src/core/SkOpts.h - diff]
src/core/SkOpts.cpp[diff]
src/core/SkUtils.cpp[diff]
src/opts/SkOpts_neon.cpp[diff]
src/opts/SkOpts_sse2.cpp[diff]
src/opts/SkUtils_opts_SSE2.cpp[Deleted - diff]
src/opts/SkUtils_opts_SSE2.h[Deleted - diff]
src/opts/SkUtils_opts_arm.cpp[Deleted - diff]
src/opts/SkUtils_opts_arm_neon.cpp[Deleted - diff]
src/opts/SkUtils_opts_none.cpp[Deleted - diff]
src/opts/opts_check_x86.cpp[diff]

14 files changed