This speaker detection via ambient sound is tough to do, and I need to know how well it works in noisier environments or during group conversations before I can have faith in it. I'd like some better control on what is stored and what isn't, as listening all day long creates issues in terms of storage. Finally, it would be nice to know more about the criteria for defining "moments."