A system and method for sequence extraction using screenshot images to generate a robotic process automation workflow is disclosed. The system and method include capturing a plurality of screenshots of steps performed by a user on an application using a processor, storing the screenshots in memory, determining action clusters from the captured screenshots by randomly clustering actions into an arbitrary predefined number of clusters, wherein screenshots of different variations of a same action is labeled in the clusters, extracting a sequence from the clusters, and discarding consequent events on the screen from the clusters, and generating an automated workflow based on the extracted sequences.