Disclosed herein are systems, methods, and computer-readable storage media for improving automatic speech recognition performance. Word lines are formed on a semiconductor substrate Bit lines are formed separated from the word lines and perpendicular to the word lines. The first lens group comprises an assembly of at least a first lens element and a second lens element, a transparent body and a third lens element which are sequentially arranged from the object side to the image plane side.