Re: [PATCH] mm: Speed up mremap on large regions

From: Kirill A. Shutemov
Date: Wed Oct 10 2018 - 06:00:21 EST

Next message: Krzysztof Kozlowski: "Re: [PATCH] config: arm: exynos: remove PROVE_LOCKING from defconfig"
Previous message: Michal Hocko: "Re: [PATCH v5 4/4] mm: Defer ZONE_DEVICE page initialization to the point where we init pgmap"
In reply to: Joel Fernandes: "Re: [PATCH] mm: Speed up mremap on large regions"
Next in thread: Joel Fernandes: "Re: [PATCH] mm: Speed up mremap on large regions"
Messages sorted by: [ date ] [ thread ] [ subject ] [ author ]

On Tue, Oct 09, 2018 at 04:04:47PM -0700, Joel Fernandes wrote:
> On Wed, Oct 10, 2018 at 01:02:22AM +0300, Kirill A. Shutemov wrote:
> > On Tue, Oct 09, 2018 at 01:14:00PM -0700, Joel Fernandes (Google) wrote:
> > > Android needs to mremap large regions of memory during memory management
> > > related operations. The mremap system call can be really slow if THP is
> > > not enabled. The bottleneck is move_page_tables, which is copying each
> > > pte at a time, and can be really slow across a large map. Turning on THP
> > > may not be a viable option, and is not for us. This patch speeds up the
> > > performance for non-THP system by copying at the PMD level when possible.
> > >
> > > The speed up is three orders of magnitude. On a 1GB mremap, the mremap
> > > completion times drops from 160-250 millesconds to 380-400 microseconds.
> > >
> > > Before:
> > > Total mremap time for 1GB data: 242321014 nanoseconds.
> > > Total mremap time for 1GB data: 196842467 nanoseconds.
> > > Total mremap time for 1GB data: 167051162 nanoseconds.
> > >
> > > After:
> > > Total mremap time for 1GB data: 385781 nanoseconds.
> > > Total mremap time for 1GB data: 388959 nanoseconds.
> > > Total mremap time for 1GB data: 402813 nanoseconds.
> > >
> > > Incase THP is enabled, the optimization is skipped. I also flush the
> > > tlb every time we do this optimization since I couldn't find a way to
> > > determine if the low-level PTEs are dirty. It is seen that the cost of
> > > doing so is not much compared the improvement, on both x86-64 and arm64.
> >
> > Okay. That's interesting.
> >
> > It makes me wounder why do we pass virtual address to pte_alloc() (and
> > pte_alloc_one() inside).
> >
> > If an arch has real requirement to tight a page table to a virtual address
> > than the optimization cannot be used as it is. Per-arch should be fine
> > for this case, I guess.
> >
> > If nobody uses the address we should just drop the argument as a
> > preparation to the patch.
>
> I couldn't find any use of the address. But I am wondering why you feel
> passing the address is something that can't be done with the optimization.
> The pte_alloc only happens if the optimization is not triggered.

Yes, I now.

My worry is that some architecture has to allocate page table differently
depending on virtual address (due to aliasing or something). Original page
table was allocated for one virtual address and moving the page table to
different spot in virtual address space may break the invariant.

> Also the clean up of the argument that you're proposing is a bit out of scope
> of this patch but yeah we could clean it up in a separate patch if needed. I
> don't feel too strongly about that. It seems cosmetic and in the future if
> the address that's passed in is needed, then the architecture can use it.

Please, do. This should be pretty mechanical change, but it will help to
make sure that none of obscure architecture will be broken by the change.

--
Kirill A. Shutemov

Next message: Krzysztof Kozlowski: "Re: [PATCH] config: arm: exynos: remove PROVE_LOCKING from defconfig"
Previous message: Michal Hocko: "Re: [PATCH v5 4/4] mm: Defer ZONE_DEVICE page initialization to the point where we init pgmap"
In reply to: Joel Fernandes: "Re: [PATCH] mm: Speed up mremap on large regions"
Next in thread: Joel Fernandes: "Re: [PATCH] mm: Speed up mremap on large regions"
Messages sorted by: [ date ] [ thread ] [ subject ] [ author ]