[LA-SiGMA-gpu] Supada Laosooksathit (candidate to lead the GPU team) will talk at tomorrow 11am
Ka Ming Tam
phy.kaming at gmail.com
Wed Jun 5 10:49:49 CDT 2013
Reminder for the talk today in 10 minutes.
On 4 June 2013 15:35, Ka Ming Tam <phy.kaming at gmail.com> wrote:
> Dear GPU Group Members,
>
> A candidate for replacing Zhifeng to lead the GPU team, Supada
> Laosooksathit,
> will give a presentation tomorrow 11am at 338 Johnston.
>
> Please see the following for the title and abstract of her talk.
>
> Best,
> Ka-Ming
>
> Time : 11am June 5th (tomorrow).
>
> Place : 338 Johnston
>
> Title : Performance Model for Predicting Run-time of a GPGPU Application
>
> Abstract:
>
> Due to the fact that the reliability of very large scaled systems is
> inversely related to the number of computing elements, fault tolerance has
> become a major concern in high performance computing (HPC) including the
> most recent deployment with GPUs. Many fault tolerance strategies, such as
> checkpoint/restart mechanism, have been studied in order to mitigate
> failures and the effectiveness.
>
> This talk provides an idea to model techniques that explore interplay
> between application performance and system reliability. More importantly,
> these two parameters play significant roles toward an optimal outcome in
> mitigating faults in a very large system. We taken into account in our
> fault tolerance strategies that balance between the application
> time-to-completion, and the time-to-failure (TTF) of the system. The former
> factor can be estimated by a performance model while the latter can be
> approximated by a reliability model.
>
On 4 June 2013 15:35, Ka Ming Tam <phy.kaming at gmail.com> wrote:
> Dear GPU Group Members,
>
> A candidate for replacing Zhifeng to lead the GPU team, Supada
> Laosooksathit,
> will give a presentation tomorrow 11am at 338 Johnston.
>
> Please see the following for the title and abstract of her talk.
>
> Best,
> Ka-Ming
>
> Time : 11am June 5th (tomorrow).
>
> Place : 338 Johnston
>
> Title : Performance Model for Predicting Run-time of a GPGPU Application
>
> Abstract:
>
> Due to the fact that the reliability of very large scaled systems is
> inversely related to the number of computing elements, fault tolerance has
> become a major concern in high performance computing (HPC) including the
> most recent deployment with GPUs. Many fault tolerance strategies, such as
> checkpoint/restart mechanism, have been studied in order to mitigate
> failures and the effectiveness.
>
> This talk provides an idea to model techniques that explore interplay
> between application performance and system reliability. More importantly,
> these two parameters play significant roles toward an optimal outcome in
> mitigating faults in a very large system. We taken into account in our
> fault tolerance strategies that balance between the application
> time-to-completion, and the time-to-failure (TTF) of the system. The former
> factor can be estimated by a performance model while the latter can be
> approximated by a reliability model.
>
-------------- next part --------------
An HTML attachment was scrubbed...
URL: https://mail.loni.org/mailman/private/lasigma-gpu/attachments/20130605/bc31e6b1/attachment.html
More information about the LASiGMA-gpu
mailing list