PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 10, 200374 citations

Finding recurrent sources in sequences

View Full Paper
AGAristides GionisHMHeikki Mannila

Key Points

Key points are not available for this paper at this time.

Abstract

Many genomic sequences and, more generally, (multivariate) time series display tremendous variability. However, often it is reasonable to assume that the sequence is actually generated by or assembled from a small number of sources, each of which might contribute several segments to the sequence. That is, there are h hidden sources such that the sequence can be written as a concatenation of k h pieces, each of which stems from one of the h sources. We define this (k, h)segmentation problem and show that it is NP-hard in the general case. We give approximation algorithms achieving approximation ratios of 3 for the L1 error measure and √ 5 for the L2 error measure, and generalize the results to higher dimensions. We give empirical results on real (chromosome 22) and artificial data showing that the methods work well in practice.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Gionis et al. (2003) studied this question.

synapsesocial.com/papers/6a17d1693275b64d0e6f28bfhttps://doi.org/10.1145/640075.640091
Ask AI
Helpful
Bookmark
Share
View Full Paper