Monday, May 5, 2014

My recent Mobile and Search Related Blog Posts

I've recently written a few blog posts at the Memkite Blog and blog.amundtveit.com

1. Privacy Efficiency - Measuring and Improving

Privacy Efficiency
Privacy Efficiency - Measuring and Improving
- About quantifying and measuring the value of privacy, inspired by how energy efficiency is measured.

2. Mobile Eats the Cloud

Mobile Eats the Cloud
Mobile Eats the Cloud
- About how powerful mobile devices are becoming (wrt storage and cpu), and how that can change things. It is written as an answer to Chris Dixon's blog post about lack of innovation in Mobile.

3. Trondheim - the Unknown Mobile Tech Capital of the Nordics


Trondheim - The Unknown Mobile Tech Capital of the Nordics
- About how Trondheim mobile tech impacts billions of people daily

Best regards, Amund

Thursday, March 29, 2012

Syllable-based forecast of best performing yc-startups from March 2012 batch

As previously written in predicting startup performance with syllables those with few syllables - typically 1 or 2 - in their name is likely to perform the best (there are of course notable exceptions, a relatively recent one being pinterest.com and instagram.com). I also wrote a prediction of the fall 2011 ycombinator batch (including rough validation of hypothesis in the first posting) 

The lists below include the current alexa rank (as a rough estimate of traffic and traction), consider that a starting point, and then growth from here can be validated later. The hypothesis is that the average growth for the top list will be significantly higher than for the bottom list. 

note to self: 
 it would be interesting to later test hypotheses around well-known english words (e.g. pair and ark) vs "functional names" (e.g. 99dresses) vs "syntetic/wordsmith names" (e.g. dealupa). Other tests would be ratio of vovels to non-vovels and impact on traffic/traction.

A-list - Predicted best performers from the ycombinator 2012 batch

Ark  - 1 syllable - 255,810 (alexa traffic rank)
Chute - 1 syllable - 3,799,189
Crowdtilt - 2 syllables - 161,411
Exec - 2 syllables - 3,623,916
Flutter - 2 syllables - 611,535 (domain: flutter.io)
Flypad - 2 syllables - 23,111,805
Kyte  - 2 syllables - 669,753
Givespark - 2 syllables - 4,900,872
Hackpad - 2 syllables - 201,595
Minefold - 2 syllables - 783,034
Midnox - 2 syllables - 3,152,778
Pair - 2 syllables - 10,045
PlanGrid - 2 syllables - 887,681
Popset - 2 syllables - 429,324 (big in japan)
Screenleap - 2 syllables -  185,485
SendHub - 2 syllables - 190,276
TiKL  - 2 syllables - 23,917,339

B-list - The rest of the startups from the batch (> 2 syllables)
42Floors  - 197,479 (alexa rank)
99dresses - 606,730
AnyVivo - no  data
Carsabi - 212,484
Coderwall - 90,456
Daily Muse - 4,320,906
Dealupa - 411,214
EveryArt - 289,105
FamilyLeaf - 606,730
HireArt - 217,183
Lvl6 - 2,612,633
Matterport - 1,749,590
Medigram - 2,427,117
Per Vices - 3,274,070
Priceonomics - 100,322
Shoptiques - 256,195
Socialcam - 46,235
Sonalight - 1,832,186
Your Mechanic - 3,650,942
Zillabyte - 1,089,060


disclaimer:
This method only counts syllables in the name and is not scientifically validated, so no reason to be offended if your startup doesn't make the A-list. There might also be erroneous counts in syllables , please let me know if you find one. 

Wednesday, August 24, 2011

Syllable-based forecast of best performing yc-startups from latest batch

As previously written in predicting startup performance with syllables those with few syllables - typically 1 or 2 - in their name is likely to perform the best. Here is a quick prediction based on the yc startups from the latest batch just published: 1 and 2-syllables (most likely to be high performers according to the few-syllable prediction)
  • MixRank (2), alexa-rank: 43,724
  • Picplum (2), alexa-rank: 700,278
  • Depteye (2), alexa-rank: no data
  • Envolve (2), alexa-rank: 52,418
  • Quartzy (2), alexa-rank: 785,102
  • Snapjoy (2), alexa-rank: 450,142
  • Opez (2), alexa-rank: 719,612
  • Stypi (2), alexa-rank: 531,965
  • ZigFu (2), alexa-rank: 21,914,936
  • Parse (1), alexa-rank: 564,829
  • Verbling (2, alexa-rank: 237,636
  • Vidyard (2), alexa-rank: 508,951
  • Tagstand (2), alexa-rank: 812,669
  • Kicksend (2, alexa-rank: 326,583
  • Can'tWait(2), alexa-rank: 786,213
I am sure the rest of the yc batch startups are just fine, but not according to syllable-based prediction, the ones I have in mind are:
  • Aisle50 (3), alexa-rank: 1,106,951
  • Launchpad Toys (3), alexa-rank: 1,425,628
  • Interviewstreet (3), alexa-rank: 135,296
  • DoubleRecall (4), alexa-rank: 799,538
  • Munch on Me (3), alexa-rank: 130,234
  • PageLever (3), alexa-rank: 55,462
  • MarketBrief (3), alexa-rank: 467,943
  • MobileWorks (3), alexa-rank: 315,150
  • Vimessa (3), alexa-rank: 314,150
  • Codeacademy (3), alexa-rank: 61,582
How did it go with the last prediction round?
Unfortunately I only added the ones I though would be best performing according to syllable-count (and not the rest for comparison) for the yc summer 2010 batch, but here is how the low syllable count  did:
  • AdGrok (acquired by Twitter)
  • Brushes (winner Apple design award 2010)
  • FanVibe (acquired by beRecruited)
  • Gantto (customers: Fujitsu, Lucasfilm++,  investor:500startups)
  • GazeHawk (investor: 500startups)
  • HipMunk (most successful startup from that yc batch?)
  • OhLife (not sure how they have done)
  • TeeVox (not sure how they have done)
Whether this beats throwing darts (random selection) is yet to be tested.
Disclaimer: If I've counted number of syllables wrong for some of the startups (have never heard pronounciation of the startup names) please ping me.

Friday, August 12, 2011

slowly back in the academic publishing game

a very long time ago I created a list of conference Call for Papers (CFP) I wanted to follow (as a fresh PhD student), this grew into a service of its own (and development shifted from me to another developer), and has been used by researchers both directly as a service and later as a part of linked data (semantic web) input, and I just became a sidekick on a poster about linked data and call for papers from that service. Hope to get the time to publish more in academic forums (in addition to blog posts) later this year, perhaps about Atbrox search technology and (forthcoming) services.

Wednesday, June 15, 2011

Mapreduce Algorithms and Search

My latest postings on the Atbrox blog:


Sunday, May 30, 2010

Evaluation of Search Predictions made in May 2000

In May 2000 I wrote A few thoughts about the future of Internet Information Retrieval (i.e. search), but how did it actually go? I've tried to evaluate them in this posting, with the original prediction in italic font followed by the evaluation.

1a) Prediction - Specialized Services within Search

It seems likely that the specialization in the Internet Information Retrieval (IIR) business will continue. Internet information crawling, pre-processing, indexing, searching and presentation requires different types of technologies and know-how, this might create opportunities for new companies specializing in only one step of the IIR "food chain". One possibility could be that companies doing crawling will do offer extracts of relevant data on request, e.g. a search engine specializing in winter sports could get only relevant data extracted from several regional crawler companies. In other words, the IIR "food chain" might increase in length.


1b) Evaluation
Specalization of search services happened to some degree, but had relatively small impact. Examples of such services include fetching/crawl-related services (e.g. 80legs). But the services with biggest impact are the free (e.g. Google Ajax Search API and Bing APIs) and commercial search APIs (e.g. Yahoo Boss and Wolfram Alpha API), all in common that they offer the last step, i.e. search - so implicitly covering all steps. Noteworthy happenings in the related direction is cloud computing and increasing number of large data sets (e.g. infochimps collection, DBPedia and the Public Terabyte (crawl) dataset)

2a) Prediction about Potential New Search Players

As the importance of Internet Information Retrieval grows, players that have been concentrating on the lower end of the Internet "food chain", i.e. major bandwidth providers (e.g. MCI or British Telecom) and network software/hardware vendors (e.g. 3COM or Cisco) might want to enter the market as providers of partially indexed data to search engines and topic hierarchies.


2b) Evaluation
This didn't happen at all to my knowledge.

3a) Prediction about Potential New Search Technologies

With the increased growth of the amount of data on the Internet, new technologies for doing distributed indexing/search of data will probably occur. This is particularly interesting if processing and indexing of multimedia data (e.g. sound, pictures and video) becomes popular. Processing of multimedia data is considerably more CPU intensive than processing of textual data. Example of such processing could be automatic detection of objects (e.g. a car) in video frames.


3b) Evaluation
(Massively) distributed indexing in the "SETI@home-style" didn't happen at large scale, though there are a few examples pursuing distributed indexing/search, e.g. the Majestic project. The in retrospective obvious processing of multimedia data is happening (but not trivial problems to solve).

Conclusion
If I am kind - 0.5 on prediction 1, 0 on prediction 2 and 0.5 on prediction 2 ~ 33.33% correct?

Sunday, January 24, 2010

My recent reads in Information Retrieval - Indexing


Information Retrieval (IR) - better known as Search - is probably the most exciting research field I know of, the reasons that makes IR exciting are:
  • solvability - it can probably never be solved perfectly, but always be improved
  • coverage - it spans all areas of computer science and touches many other sciences (e.g. statistics)
  • importance - it is the most important research area related to supporting human decisions? (~AI)
  • difficulty - it is extremely hard to do well
  • applicability - it can be used practically anywhere (anytime).
Where to start learning about information retrieval?
Before jumping into research papers I suggest reading a book about IR, either:
Search Engines: Information Retrieval in Practice (2009) or
Introduction to Information Retrieval (2008)
They are both good and relatively similar books written by a mix of authors from search industry and academic IR research (note: I personally prefer the newest one).

My recent reads in Information Retrieval?


Indexing - algorithms and datastructures for self-indexing
Self-indexing is where (lossless) compression meets indexing, and is an alternative to the classic inverted index. Self-indices has some nice characteristics wrt compression, performance and query-flexibility. Indexing-research-rockstar Gonzalo Navarro even called it the Miracle of Self-indexing (2009).
2 key papers in the field are:
  1. Opportunistic Data Structures with Applications (2000)
    • Introduced the FM-index
  2. High-order entropy-compressed text indexes (2003)
    • Introduced the Wavelet Index Tree
Check out Navarro's survey paper Compressed Full-Text Indexes (2007) for a good overview of self-indexing.

Have a nice read :)

Wednesday, October 7, 2009

Sunday, March 8, 2009

Snakes on a Cloud

.. or Mobile Agents with Python

What are mobile agents?
"A mobile agent is a process that can transport its state from one environment to another, with its data intact, and be capable of performing appropriately in the new environment", source: Wikipedia.

Why Mobile Agents .. or how to deal with Data Gravity
The phrase "with its data" most likely refers to the agent's state data (and not all types of data), since one of the nicest properties of mobile agents is that they can move to where potentially huge amounts of data is. By plotting the curve "data gravity" - i.e. a rough estimate of how long time it takes to empty/fill/process a hard drive over a network connection - hard disk size divided by (typical) ethernet network speed over the last 20 years - the motivation for moving code to data (and not vice versa) is clearly increasing, making mobile agents a potentially interesting approach.

Mobile Agent Runtime Environment
A basic requirement for a mobile agent runtime environment is the ability to receive and run the agent's code, e.g. typically support one/several of the following (with Python-related examples):
i) receive and run binary code
python example: receive python compiled to binary with Shedskin and g++
ii) receive and compile source code and run it
python example: receive c source code and compile/integrate it with Python using Cinpy, or use Shedskin/g++ on received Python code
iii) receive and run interpreter on source code
python example: receive python code and interpret using the eval() method.

Bandwidth
If the mobile agents move around a bit, you probably want their representation as compact as possible to reduce bandwidth requirements, i.e. prefer agents represented in (small amounts of) source code - alternative ii) or iii) - over (larger amounts of) binary code - alternative i).

Processing
compiled code (or just-in-time compiled code) is usually more efficient than interpreted code (interpreted code can perhaps be seen as analog to the "gas guzzling cars" of computing wrt resource utilization, fortunately there are tools to deal with that), so alternative ii) is probably preferred over iii), and with the cinpy case compilation overhead is negligible (a few milliseconds to compile and make a short C method ready to be called from Python), which matter if you have a large amount of distributed mobile agents. Pareto principle also matters for mobile agents, so a mix (80-99% of code) in interpreted Python and the things that really need to perform in C (quickly compiled with cinpy) might be a common mix.

Example of Cinpy-wrapped C function in python
fibc=cinpy.defc(
"fib",
ctypes.CFUNCTYPE(ctypes.c_int,ctypes.c_int),
"""
int fib(int x) {
if (x<=1) return 1; return fib(x-1)+fib(x-2); } """)


Security
In case mobile agents move around on the cloud it is nice to know that the agent you receive is from a known source (yourself), this can e.g. be done using ezPyCrypto's signString() to sign the agent source and then use verifyString() methods on the signed agent source code together with its signature to check the origin of the agent (assuming the signer's public key is available on the receiving end).

disclaimer: this posting (and all others on this blog) only represents my personal views.

Saturday, December 27, 2008

Ajax with Python - Combining PyJS and Appengine with json-rpc

I recently (re)discovered pyjs  - also called Pyjamas - which  is a tool to support development of (client-side) Ajax applications with Python, it does that by compiling Python code to Javascript (Pyjamas is inspired by GWT  - which supports writing Ajax applications in Java).

pyjs' kitchensink comes with a JSON-RPC example, this posting shows how to use Appengine to serve as a JSON-RPC server for the pyjs json-rpc example.

(Rough) Steps:
1) download appengine SDK and create an appengine application (in the dashboard)
2) download pyjs 
3) Replace pyjs example/jsonrpc/JSONRPCExample.py with this code

from ui import RootPanel, TextArea, Label, Button, HTML, VerticalPanel, HorizontalPanel, ListBox


from JSONService import JSONProxy


class JSONRPCExample:
    def onModuleLoad(self):
        self.TEXT_WAITING = "Waiting for response..."
        self.TEXT_ERROR = "Server Error"
        self.remote_py = UpperServicePython()
        self.status=Label()
        self.text_area = TextArea()
        self.text_area.setText(r"Please uppercase this string")
        self.text_area.setCharacterWidth(80)
        self.text_area.setVisibleLines(8)
        self.button_py = Button("Send to Python Service", self)
        buttons = HorizontalPanel()
        buttons.add(self.button_py)
        buttons.setSpacing(8)
        info = r'This example demonstrates the calling of appengine upper(case) method with JSON-RPC from javascript (i.e. Python code compiled with pyjs to javascript).'
        panel = VerticalPanel()
        panel.add(HTML(info))
        panel.add(self.text_area)
        panel.add(buttons)
        panel.add(self.status)
        RootPanel().add(panel)


    def onClick(self, sender):
        self.status.setText(self.TEXT_WAITING)
        text = self.text_area.getText()
        if self.remote_py.upper(self.text_area.getText(), self) < 0:
            self.status.setText(self.TEXT_ERROR)


    def onRemoteResponse(self, response, request_info):
        self.status.setText(response)


    def onRemoteError(self, code, message, request_info):
        self.status.setText("Server Error or Invalid Response: ERROR " + code + " - " + message)


class UpperServicePython(JSONProxy):
    def __init__(self):
        JSONProxy.__init__(self, "/json", ["upper"])



4) Use the following code for the appengine app
from google.appengine.ext import webapp
from google.appengine.ext.webapp.util import run_wsgi_app
import logging
from django.utils import simplejson

class JSONHandler(webapp.RequestHandler):
  def json_upper(self,args):
    return [args[0].upper()]

  def post(self):
    args = simplejson.loads(self.request.body)
    json_func = getattr(self, 'json_%s' % args[u"method"])
    json_params = args[u"params"]
    json_method_id = args[u"id"]
    result = json_func(json_params)
    # reuse args to send result back
    args.pop(u"method")
    args["result"] = result[0]
    args["error"] = None # IMPORTANT!!
    self.response.headers['Content-Type'] = 'application/json'
    self.response.set_status(200)
    self.response.out.write(simplejson.dumps(args))

application = webapp.WSGIApplication(
                                     [('/json', JSONHandler)],
                                     debug=True)

def main():
  run_wsgi_app(application)

if __name__ == "__main__":
  main()
5) compile pyjs code in 3) and create static dir in appengine app to store compiled code (i.e in app.yaml)
6) use dev_appserver.py to test locally or appcfg.py to deploy on appengine

Facebook application
By using this recipe it was easy to create a facebook app of the example in this posting - check it out at http://apps.facebook.com/testulf/


Alternative backend - using webpy


This shows how to use webpy as backend, just put the javascript/html resulting from pyjs compile into the static/ directory. Very similar as the previous approach, with the code in blue being the main diff (i.e. webpy specific)
import web
import json
urls = (
'/json', 'jsonhandler'
)
app = web.application(urls, globals())

class jsonhandler:
  def json_upper(self,args):
    return [args[0].upper()]

  def json_markdown(self,args):
    return [args[0].lower()]

  def POST(self):
    args = json.loads(web.data())
    json_func = getattr(self, 'json_%s' % args[u"method"])
    json_params = args[u"params"]
    json_method_id = args[u"id"]
    result = json_func(json_params)
    # reuse args to send result back
    args.pop(u"method")
    args["result"] = result[0]
    args["error"] = None # IMPORTANT!!
    web.header("Content-Type","text/html; charset=utf-8")
    return json.dumps(args)

if __name__ == "__main__":
  app.run()

Wednesday, December 17, 2008

Simulating Mamma Mia under the xmas tree (with Python)

Norway has a population of ~4.7 million, and "Mamma Mia!" has during a few weeks been sold in ~600 thousand DVD/Blueray copies (> 12% of the population, i.e. breaking every sales record, perhaps with the exception of Sissel Kyrkjebø's Christmas Carol album which has sold a total of ~900 thousand copies in the same population, but that was over a period of 21 years).  In the UK it has sold more than 5 million (in a population of 59 million, i.e. >8% of the population).

Well, to the point. Guesstimating that there will be ~2 million xmas trees in Norway, one can assume that many of the trees will have (much) more than one "Mamma Mia!" dvd/blueray underneath it*,  the question is how many?

Simulating Mamma Mia under the xmas tree
.. Before you start recapping probability theory, combinatorics, birthday paradox, multiplication of improbabilities and whatnot, how about finding an alternative solution using simulation

A simple simulation model could be to assume that for every Mamma Mia dvd/blueray copy there is a lottery where all trees participate,  and then finally (or incrementally) count how many copies each tree won.

A friend wrote a nice Python script to simulate this:

import random

NUM_XMAS_TREES = 2000000
NUM_MAMMA_MIAS = 600000

tree_supplies = {}
for mamma_mia in range(NUM_MAMMA_MIAS):
    winner_tree = random.randint(0, NUM_XMAS_TREES)
    tree_supplies[winner_tree] = tree_supplies.get(winner_tree,0) + 1

tree_stats = {}
for tree in tree_supplies:
    tree_stats[tree_supplies[tree]] = tree_stats.get(tree_supplies[tree], 0) + 1

print tree_stats

Results:

$ for k in `seq 1 10`; do echo -n "$k " ; python mammamia.py; done
1 {1: 443564, 2: 67181, 3: 6618, 4: 510, 5: 36}
2 {1: 444497, 2: 66811, 3: 6543, 4: 520, 5: 32, 6: 2}
3 {1: 444796, 2: 66376, 3: 6738, 4: 510, 5: 36, 6: 3}
4 {1: 444499, 2: 66750, 3: 6652, 4: 469, 5: 30, 6: 2, 7: 1}
5 {1: 444347, 2: 66717, 3: 6697, 4: 494, 5: 28, 6: 2}
6 {1: 444511, 2: 66389, 3: 6763, 4: 551, 5: 40, 6: 3}
7 {1: 443914, 2: 66755, 3: 6785, 4: 511, 5: 33, 6: 2}
8 {1: 444747, 2: 66558, 3: 6667, 4: 484, 5: 40}
9 {1: 444553, 2: 66703, 3: 6631, 4: 497, 5: 32}
10 {1: 443903, 2: 66853, 3: 6774, 4: 487, 5: 23, 6: 1}

Conclusion

So we see that in run 4 there was one xmas tree that according to the simulation model got 7(!) Mamma Mia DVD/Bluerays underneath it, but the overall simulation shows that 5 or 6 (at most) is probably more likely (assuming the model is right).

Regarding the simulation model, it is probably way too simplistic, i.e. not taking into account people buying mamma mia for themselves (or not as xmas gifts), a likely skewness in terms of number of gifts per christmas tree, interaction between buyers, etc. But it can with relatively simple manners be extended with code to make it more realistic. Check out Markov Chain Monte Carlo simulation for more info on how to create more realistic simulation models.

Wednesday, December 3, 2008

cinpy - or C in Python

cinpy is a tool where you can write C code in your Python code (with the help of ctypes - included in modern Python versions). When you execute your python program the C code is compiled on the fly using Tiny C Compiler (TCC).
In this posting I will describe:
1) installation and testing cinpy
2) a simple benchmark (c-in-py vs python)
3) compare performance with gcc (c-in-py vs gcc)
4) measure cinpy (on-the-fly) compilation time
5) how to dynamically change cinpy methods

1. How to install and try cinpy (note: also found in cinpy/tcc README files)
1) download, uncompress and compile TCC
./configure
make
make install
gcc -shared -Wl,-soname,libtcc.so -o libtcc.so libtcc.o

2) download, uncompress and try cinpy
cp ../tcc*/*.so .
python cinpy_test.py  # you may have to comment out or install psyco 

2. Sample performance results (on a x86 linux box):

python cinpy_test.py
Calculating fib(30)...
fibc : 1346269 time: 0.03495 s
fibpy: 1346269 time: 2.27871 s
Calculating for(1000000)...
forc : 1000000 time: 0.00342 s
forpy: 1000000 time: 0.32119 s

Using cinpy for fibc (Fibonacci) method was ~65 times faster than fibpy, and and cinpy for forc (loop) was ~93 times faster than forpy, not bad.

3. How does cinpy (compiled with tcc) compare to gcc performance?
Copying the C fib() method and calling it with main program
$ time fibgcc


fib(30) = 1346269

real    0m0.016s
user    0m0.020s
sys     0m0.000s


GCC gives roughly twice as fast code as cinpy/tcc (0.034/0.016). 




4. How long time does it take for tcc to on-the-fly compile cinpy methods?


#!/usr/bin/env python
# cinpy_compile_performance.py
import ctypes
import cinpy
import time
t=time.time
t0 = t()
fibc=cinpy.defc(
     "fib",
     ctypes.CFUNCTYPE(ctypes.c_int,ctypes.c_int),
     """
     int fib(int x) {
       if (x<=1) return 1;
       return fib(x-1)+fib(x-2);
     }
     """)
t1 = t()
print "Calculating fib(30)..."
sc,rv_fibc,ec=t(),fibc(30),t()
print "fibc :",rv_fibc,"time: %6.5f s" % (ec-sc)
print "compilation time = %6.5f s" % (t1-t0)

python cinpy_compile_performance.py

Calculating fib(30)...
fibc: 1346269 time: 0.03346 s
compilation time: 0.00333 s


So compilation (and linking) time is about 3.3ms, which is reasonably good, not a lot of overhead! (note: TCC is benchmarked to compile, assemble and link 67MB of code in 2.27s)


5. "Hot-swap" replacement of cinpy code?
Let us assume you have a system with a "pareto" situation, i.e. 80-99% of the code doesn't need be fast (written in Python), but 1-20% need to be really high performance (and written in C using cinpy), and that you need to frequently change the high performing code, can that be done? Examples of such a system could be mobile (software) agents.


Sure, all you need to do is wrap your cinpy definition as a string and run exec on it, like this:

origmethod="""
fibc=cinpy.defc("fib",ctypes.CFUNCTYPE(ctypes.c_int,ctypes.c_int),
          '''
          int fib(int x) {
              if (x<=1) return 1;
              return fib(x-1)+fib(x-2);
          }
          ''')
"""
# add an offset to the Fibonacci method
alternatemethod = origmethod.replace("+", "+998+")
print alternatemethod
# alternatemethod has replaces origmethod with exec(alternatemethod)
print fibc(2) # = 1000
Conclusion:
cinpy ain't bad.

Remark: depending on the problem you are solving (e.g. if it is primarily network IO bound and not CPU bound) , becoming 1-2 orders of magnitude faster (cinpy vs pure python) is probably fast enough (CPU wise), the doubling from GCC may not matter (since network IO wise Python performs quite alright, e.g. with Twisted ).