On Optimal Multiple Changepoint Algorithms for Large Data

Maidstone, Robert; Hocking, Toby; Rigaill, Guillem; Fearnhead, Paul

统计学 > 方法论

arXiv:1409.1842 (stat)

[提交于 2014年9月5日 ]

标题：关于大样本数据的最优多重变点算法

标题： On Optimal Multiple Changepoint Algorithms for Large Data

Authors:Robert Maidstone, Toby Hocking, Guillem Rigaill, Paul Fearnhead

摘要：对于能够准确检测长时间序列或多等效数据中的变化点的需求日益增加。许多常见的变化点检测方法（例如基于惩罚似然或最小描述长度的方法）都可以表述为最小化分段的成本函数。存在精确解决此最小化问题的动态规划方法，但这些方法通常在时间序列长度上至少呈二次增长。存在计算成本接近线性的时间序列长度的算法（例如二分分割），但它们无法保证找到最优分段。最近提出了加速动态规划算法的想法，同时仍能保证找到成本函数的真实最小值。在这里，我们扩展了这些剪枝方法，并引入了两种用于分割数据的新算法：FPOP 和 SNIP。经验结果显示，FPOP 比现有的动态规划方法快得多，并且与现有方法不同，其计算效率不受数据中变化点数量的影响。我们评估了该方法在检测拷贝数变异方面的性能，并观察到 FPOP 的计算成本与二分分割法具有竞争力。

摘要： There is an increasing need for algorithms that can accurately detect changepoints in long time-series, or equivalent, data. Many common approaches to detecting changepoints, for example based on penalised likelihood or minimum description length, can be formulated in terms of minimising a cost over segmentations. Dynamic programming methods exist to solve this minimisation problem exactly, but these tend to scale at least quadratically in the length of the time-series. Algorithms, such as Binary Segmentation, exist that have a computational cost that is close to linear in the length of the time-series, but these are not guaranteed to find the optimal segmentation. Recently pruning ideas have been suggested that can speed up the dynamic programming algorithms, whilst still being guaranteed to find true minimum of the cost function. Here we extend these pruning methods, and introduce two new algorithms for segmenting data, FPOP and SNIP. Empirical results show that FPOP is substantially faster than existing dynamic programming methods, and unlike the existing methods its computational efficiency is robust to the number of changepoints in the data. We evaluate the method at detecting Copy Number Variations and observe that FPOP has a computational cost that is competitive with that of Binary Segmentation.

评论：	20页
主题：	方法论 (stat.ME) ; 计算 (stat.CO)
MSC 类：	62M10
引用方式：	arXiv:1409.1842 [stat.ME]
	(或者 arXiv:1409.1842v1 [stat.ME] 对于此版本)
	https://doi.org/10.48550/arXiv.1409.1842

提交历史

来自： Robert Maidstone [查看电子邮件]
[v1] 星期五， 2014 年 9 月 5 日 15:44:34 UTC (307 KB)

统计学 > 方法论

标题：关于大样本数据的最优多重变点算法

标题： On Optimal Multiple Changepoint Algorithms for Large Data

提交历史

获取论文：

参考文献与引用

收藏

文献和引用工具

与本文相关的代码，数据和媒体

演示

推荐器和搜索工具

arXivLabs：与社区合作伙伴的实验项目

统计学 > 方法论

标题： 关于大样本数据的最优多重变点算法 显示英文标题

标题： On Optimal Multiple Changepoint Algorithms for Large Data

提交历史

获取论文：

参考文献与引用

BibTeX 格式的引用

收藏

文献和引用工具

与本文相关的代码，数据和媒体

演示

推荐器和搜索工具

arXivLabs：与社区合作伙伴的实验项目

标题：关于大样本数据的最优多重变点算法