It does work. With stacks of pre-trained sparse convolutional autoencoders, we can achieve a strong starting point for fine-tuning on small labeled datasets (like Caltech 101, which had 30 training samples per category) and reach near-state-of-the-art performance.