Parallelizing dense and banded linear algebra libraries using SMPSs

Badía Sala, Rosa María; Herrero, Josep R.; Labarta Mancho, Jesús; Pérez, Josep M.; Quintana-Orti, Enrique S.; Quintana-Ortí, Gregorio

dc.contributor.author	Badía Sala, Rosa María
dc.contributor.author	Herrero, Josep R.
dc.contributor.author	Labarta Mancho, Jesús
dc.contributor.author	Pérez, Josep M.
dc.contributor.author	Quintana-Orti, Enrique S.
dc.contributor.author	Quintana-Ortí, Gregorio
dc.date.accessioned	2011-05-30T08:49:24Z
dc.date.available	2011-05-30T08:49:24Z
dc.date.issued	2009-12-25
dc.identifier.citation	BADIA, Rosa M., et al. Parallelizing dense and banded linear algebra libraries using SMPSs. Concurrency and Computation: Practice and Experience, 2009, vol. 21, no 18, p. 2438-2456.
dc.identifier.issn	1532-0626
dc.identifier.uri	http://hdl.handle.net/10234/22715
dc.description.abstract	The promise of future many-core processors, with hundreds of threads running concurrently, has led the developers of linear algebra libraries to rethink their design in order to extract more parallelism, further exploit data locality, attain better load balance, and pay careful attention to the critical path of computation. In this paper we describe how existing serial libraries such as (C)LAPACK and FLAME can be easily parallelized using the SMPSs tools, consisting of a few OpenMP-like pragmas and a runtime system. In the LAPACK case, this usually requires the development of blocked algorithms for simple BLAS-level operations, which expose concurrency at a finer grain. For better performance, our experimental results indicate that column-major order, as employed by this library, needs to be abandoned in benefit of a block data layout. This will require a deeper rewrite of LAPACK or, alternatively, a dynamic conversion of the storage pattern at run-time. The parallelization of FLAME routines using SMPSs is simpler as this library includes blocked algorithms (or algorithms-by-blocks in the FLAME argot) for most operations and storage-by-blocks (or block data layout) is already in place
dc.format.extent	14 p.
dc.language.iso	eng
dc.publisher	John Wiley
dc.relation.isFormatOf	Versió post-print del document publicat a: http://www3.interscience.wiley.com/journal/122519419/abstract
dc.relation.isPartOf	Concurrency and computation : practice & experience, 2009 vol. 21, no. 18
dc.rights.uri	http://rightsstatements.org/vocab/CNE/1.0/	*
dc.subject	Programmability
dc.subject	High performance
dc.subject	Dynamic scheduling
dc.subject	Multi-core processors
dc.subject.lcsh	Embedded computer systems
dc.subject.other	Sistemes incrustats (Informàtica)
dc.title	Parallelizing dense and banded linear algebra libraries using SMPSs
dc.type	info:eu-repo/semantics/article
dc.rights.accessRights	info:eu-repo/semantics/openAccess

Ficheros en el ítem

Nombre:: 34981.pdf
Tamaño:: 213.3Kb
Formato:: PDF
Descripción:: versió post-print

Ver/Abrir

Este ítem aparece en la(s) siguiente(s) colección(ones)

ICC_Articles [424]

Mostrar el registro sencillo del ítem